The growing use of copyrighted books and other published content to train AI models is creating a difficult legal debate. Companies developing systems such as ChatGPT, Gemini and Claude have trained models using enormous amounts of written material, but the question of whether that use is legally allowed remains unsettled.

The answer largely depends on how the copyrighted material was obtained, how it was used and whether the AI system’s use is considered fair use.

Why AI training and copyright are being debated

AI models learn from huge collections of text and other data. That can include books, articles, research papers and publicly available online content.

Authors have raised concerns that their work may have been used to develop AI systems without permission or payment. AI companies, meanwhile, have argued that training a model is different from simply reproducing a copyrighted book.

That distinction is becoming increasingly important in court cases.

Court decisions are not giving a simple answer

A major case involving Anthropic highlighted the complexity of the issue. A US judge ruled that using books to train AI could be lawful, but the company still faced a $1.5 billion settlement related to obtaining books from unauthorized sources.

The decision suggests that the source of training data can matter just as much as the eventual use of that data.

Another case involving Thomson Reuters and Ross Intelligence reached a different conclusion. The court found that using Thomson Reuters’ copyrighted material to build a competing AI-powered legal product did not qualify as fair use.

These cases show why there is currently no single rule that makes all AI training either legal or illegal.

Fair use could become a key factor

In the US, fair use allows certain uses of copyrighted material without permission. Courts generally consider factors such as the purpose of the use, how much copyrighted material was used and whether the use affects the original work’s market.

AI training is difficult to fit into these traditional categories because the technology is relatively new while much of the copyright framework was created decades ago.

Legal experts also point out that training an AI model and creating content with AI are two different copyright questions.

What happens next?

Most major AI copyright disputes are still being challenged through the courts. Future decisions could influence how AI companies collect training data, how authors protect their work and whether licensing becomes a more common part of AI development.

For now, businesses, developers and creators should avoid assuming that all AI training on copyrighted material is automatically legal or illegal. The legal position depends heavily on the circumstances and the jurisdiction involved.

As AI technology continues to develop, courts and lawmakers will likely play a major role in defining where copyright protection ends and AI innovation begins.

Read More on VitalStack

Enjoyed this article?

Subscribe for weekly deep-dives on AI and health — straight to your inbox.