
AI companies that train chatbots on published books are finding that copyright law, written long before anyone imagined generative AI, offers no clear answer. The models behind ChatGPT, Gemini, and Claude learn from databases containing hundreds of millions of books, articles, and academic papers — often without the authors’ knowledge or consent. Whether that practice is legal has become one of the most consequential questions in technology, and the courts are still working it out.
A $1.5 billion ruling that wasn’t what it seemed
Last year, Judge William Alsup ordered Anthropic to pay $1.5 billion to a group of writers whose works were used to train the company’s AI models. At first glance, it looked like a victory for authors. But the ruling was more complicated than the headline suggested.
Alsup actually found that the AI training itself was lawful. The penalty came from a different problem: the company had pirated books from illegal online shadow libraries to build its training data. The judge made that distinction explicitly.
Related: Pixel 11 Pro XL better camera, same basics
“Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” Alsup wrote, comparing how an AI model processes trillions of words to how a human writer studies literature.
Why fair use is at the center of the fight
Copyright law hasn’t been meaningfully updated since 1976. That leaves judges interpreting fifty-year-old guidelines for questions that could define the future of the AI industry. The core issue usually comes down to fair use — whether training on copyrighted material counts as “transformative” enough to be legal.
Fair use is a legal carve-out that permits using copyrighted works without permission in certain contexts, like criticism, parody, education, or commentary. Judges weigh several factors, including the purpose of the use, how much of the work is used, and whether it hurts the original work’s market.
Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, said courts are struggling to apply these tests consistently. “Everybody is very worried right now because the law is all over the place, and it’s because of this question,” he told TechCrunch. “They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.”
Related: The Silent Revolution: How Singapore is Redefining the Global Tech Talent Landscape
One case has given some shape to the debate. Thomson Reuters sued Ross Intelligence, a research firm, for copying its legal content to build a competing AI-based platform. Judge Stephanos Bibas ruled against Ross last year, writing that “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s.” The key factor was competition — Ross built a product that directly rivaled the source material.
Cathy Gellis, an attorney specializing in intellectual property and technology law, sees that distinction as important. “Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work,” she said. That reasoning, she believes, generally favors AI companies — provided they aren’t stealing the books in the first place.
The practical result for authors is uncomfortable. A $1.5 billion fine sounds substantial, but for a company projecting around $200 billion in annual revenue by 2028, it’s closer to a cost of doing business than a deterrent. The ruling establishes that training on copyrighted works can be legal, which effectively green-lights the core practice while punishing the piracy used to obtain some of the source material.
The separate question of AI-generated content
Gellis also points out that the debate over training AI on copyrighted works is entirely different from the question of whether AI-generated content can itself be copyrighted. In Thaler v. Perlmutter, a court ruled that a work created entirely by AI is not copyrightable. That raises its own set of problems — namely, how to prove whether a work was AI-assisted, and what percentage of AI involvement changes the legal status.
Related: WHY LOW-LOSS COAXIAL CABLE MATTERS IN WIRELESS COMMUNICATION SYSTEMS
“If you write your novel in Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel,” Gellis said. “AI is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.”
Most AI companies remain in active litigation, and no definitive resolution is coming soon. The early rulings are shaping behavior, though. “What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later states of litigation to figure out which one will prevail,” Gellis said. “But in the meantime, all these decisions are shaping everything that’s happening. It would be kind of foolish for the AI companies to ignore them.”
For now, the legal situation remains fragmented. Different courts are reaching different conclusions based on the specific facts of each case — whether the AI tool competes directly with the source, whether the training data was obtained legally, and how transformative the resulting technology actually is. Authors and AI companies alike are waiting for higher courts to resolve the contradictions, but that process could take years.


