24 Aug 2026, Mon

The Murky Legal Battlefield: How AI Training on Copyrighted Works is Reshaping Intellectual Property Law

The artificial intelligence revolution, powered by sophisticated models like ChatGPT, Gemini, and Claude, has ushered in an era of unprecedented technological advancement. These powerful chatbots, capable of generating human-like text, creative content, and complex analyses, are built upon a foundation of seemingly inexhaustible digital libraries. These vast training datasets comprise hundreds of millions of books, countless online articles, academic papers, and a substantial portion of the internet’s publicly available information. This ubiquity of data raises a critical question: have most published authors, without their explicit knowledge or consent, inadvertently contributed to the very AI tools that now pose a significant threat to their creative livelihoods? While this scenario might appear inherently illegal, the legal landscape surrounding AI training and copyright is far more intricate and contentious than a simple violation.

"I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on," Cathy Gellis, an attorney with deep expertise in intellectual property, copyright, and technology, explained to TechCrunch. "It’s very complex and there are a lot of raw feelings about what is happening, both for and against." This complexity is further amplified by the fact that copyright law, as it stands today, has not undergone significant updates since 1976. This legislative inertia forces judges to grapple with interpreting decades-old statutes in the face of cutting-edge technological developments, creating a fertile ground for legal uncertainty and divergent interpretations that are crucial in shaping the future trajectory of the AI industry.

A landmark case that offered a glimpse into this evolving legal battleground involved AI company Anthropic. In a ruling that sent ripples through the tech and publishing worlds, Judge William Alsup ordered Anthropic to pay a staggering $1.5 billion copyright settlement to a group of authors whose works had been incorporated into the training data for the company’s AI models. On the surface, this settlement appeared to be a resounding victory for authors, a clear affirmation of their intellectual property rights. However, the nuances of Judge Alsup’s decision revealed a more complicated reality. The judge, in fact, ruled that Anthropic’s use of copyrighted material for AI training was lawful. The substantial penalty was not for the act of training itself, but for the method through which Anthropic acquired these works – specifically, by pirating them from illegal online shadow libraries. This distinction is crucial: the illegal act was the unauthorized acquisition of copyrighted material, not its subsequent use in an AI training process that was deemed permissible under certain legal interpretations.

Judge Alsup’s reasoning drew an analogy between the way an AI model ingests vast quantities of text and the process by which a human writer studies literature. "Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them – but to turn a hard corner and create something different," the judge articulated. This comparison highlights a core debate within copyright law: the distinction between "copying" and "using" or "consuming" a work.

However, intellectual property experts like Cathy Gellis see this ruling as potentially more beneficial to AI companies than to authors. The $1.5 billion fine, while substantial in absolute terms, pales in comparison to Anthropic’s projected financial growth. With the company forecasting approximately $200 billion in annual revenue by 2028, the settlement might be viewed as a manageable cost of doing business rather than a prohibitive deterrent. "I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work," Gellis observed. She further elaborated on the foundational principles of copyright law: "Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work." This statement underscores the critical legal interpretation that the act of AI training, akin to reading, might not constitute copyright infringement if it doesn’t involve direct reproduction or distribution of the original works.

The legal uncertainty surrounding AI and copyright is further exacerbated by the fact that current legal frameworks are struggling to keep pace with technological advancements. "Everybody is very worried right now because the law is all over the place, and it’s because of this question," stated Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International. "They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question." This sentiment is echoed by legal scholars and practitioners alike, who recognize the urgent need for updated legislation or clearer judicial interpretations to address the unique challenges posed by AI.

At the heart of many of these legal disputes lies the doctrine of "fair use," a critical carve-out within copyright law. Fair use permits the limited use of copyrighted materials without explicit permission from the copyright holder, provided that such use is considered "transformative" enough to be legally permissible. This doctrine is designed to foster creativity, innovation, and the public discourse by allowing for commentary, criticism, parody, education, and other forms of derivative works. When determining fair use, courts typically consider several factors: the purpose and character of the use, the nature of the copyrighted work, the amount and substantiality of the portion used in relation to the copyrighted work as a whole, and the effect of the use upon the potential market for or value of the copyrighted work.

Henderson emphasizes the fundamental purpose of copyright law: "Copyright is always about protecting and growing the market." He further notes the inconsistent reasoning observed in judicial decisions regarding AI cases: "The courts are kind of all over the place in their reasoning [in AI cases]. What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay." This suggests a key differentiator in judicial interpretation: whether the AI’s training data is used to create a directly competing product or service.

Henderson’s observation is exemplified by a case involving Thomson Reuters, a prominent media and technology company, and Ross Intelligence, a research firm. Thomson Reuters sued Ross Intelligence, alleging that the latter had copied its content to build a competing, AI-based legal platform. In this instance, Judge Stephanos Bibas ruled that Ross Intelligence’s use was not transformative. He reasoned that it did not possess a "further purpose or different character" from that of Thomson Reuters’ original content. Consequently, the court deemed it not to be fair use to train an AI on Reuters’ content for the explicit purpose of creating a direct competitor. While authors might argue that chatbots, by generating new synthetic books based on their works, are indeed competing with them, this specific argument has yet to gain widespread traction and prevail definitively in the courts.

Cathy Gellis suggests that a clearer understanding of the AI and copyright relationship can be achieved by distinguishing between two distinct, yet related, issues: the copyright implications of AI training data and the copyrightability of AI-generated content. These are fundamentally different legal questions, each with its own set of challenges.

In the realm of AI-generated content, a notable case is Thaler v. Perlmutter. Here, the court ruled that a work entirely generated by AI is not copyrightable. This decision opens a Pandora’s Box of new questions, particularly concerning the ability to definitively prove whether a work was created using AI and, if so, to what extent. The line between human authorship aided by AI and purely AI-generated content becomes increasingly blurred. Gellis uses a relatable analogy: "If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel." However, she points out, AI’s pervasive involvement in the creative process is forcing a re-evaluation of long-ignored assumptions. "AI is forcing us to look at a whole bunch of decisions that we kind of ignored for a while." This necessitates a deeper examination of authorship, originality, and the role of technology in the creative process.

The legal landscape surrounding AI and copyright remains in flux, with most AI companies currently embroiled in ongoing litigation. This means that definitive solutions to these complex issues are unlikely to emerge in the immediate future. "What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later states of litigation to figure out which one will prevail," Gellis explains. Nevertheless, she cautions that these early judicial decisions are already exerting a significant influence on the development and deployment of AI technologies. "But in the meantime, all these decisions are shaping everything that’s happening. It would be kind of foolish for the AI companies to ignore them." As the legal system grapples with the profound implications of AI, the ongoing battles in courtrooms will undoubtedly continue to shape the future of both artificial intelligence and the enduring principles of intellectual property law.

Leave a Reply

Your email address will not be published. Required fields are marked *