Fair Use in AI: Court Rules in Favor of Anthropic

Long awaited, on June 23, 2025, the first historic decision on fair use was issued, marking a point in favor of Anthropic, followed two days later by a second decision that declared Meta’s victory on the same topic.

In both cases, these are interim decisions that will have to be followed by a final ruling, but this does not lessen the importance of the affirmed statements regarding the lawfulness or not of the use of copyrighted data for machine training.

Fair Use

Let us recall that fair use is a U.S. doctrine developed by jurisprudence and introduced in art. 17 of the U.S. Code, paragraph 107, which states that it is lawful to use copyrighted material without seeking the consent of the rights holders, when certain conditions are met.

Specifically, a case-by-case overall evaluation must be made of four elements:

  1. the possible transformative use of the copy,
  2. the nature of the copied work,
  3. the amount of material copied,
  4. the impact on the market of said copy.

Generally, when a transformative use is found—which occurs whenever the purpose of the copier is completely different from that for which it was created—fair use is recognized.

For example, copying a text in order to create a system that facilitates word search within it may be considered lawful, provided that the person performing this operation does not make the full text available on the market.

The Author Guild v. Google Case

This is what was decided in 2015 in the The Author Guild v. Google case concerning the Google Books service, for which Google had copied entire books protected by copyright to develop a system that facilitated internal search for certain words.

Google Books allowed users to view only small portions of text to provide context for the search results, but never the complete book.

Considering all four fair use factors, the Court found that Google’s full digital copy of the works to offer public search functionality and snippet display fell within fair use and did not violate the authors’ copyright, because it did not render purchase of the book useless, since it did not make the book available to readers, and especially because the copy was instrumental in providing a service completely different from mere reading of the text.

The Bartz et al. v. Anthropic Case

This decision is the basis of Anthropic’s defense in the first case we will focus on (C 24-05417), filed in August 2024 by a group of authors who challenged the use of their works during training, but not the outputs generated by the system.

Feeling confident in its position, Anthropic asked the judge to issue a preliminary decision solely on fair use, in order to establish from the outset whether the use of protected data could be considered lawful based on the American doctrine.

Anthropic not only cited the The Author Guild v. Google case but even hired Tom Turvey, former head of partnerships for the Google Books project, whose task was to obtain licenses for “all the books in the world”—something he did not do—opting instead for the purchase of printed copies intended to constitute the “research library” of the AI system, which risks being the banana peel that could cause them to slip.

According to court documents, the company chose to scan millions of books, then extract from them the data and information necessary for training, essentially using a technique similar to that used by Google Books, building its collection and relying on a similar legal solution in its favor.

The Court firmly ruled that using protected works to train artificial intelligence systems falls within the principle of fair use and must be considered lawful.

The purpose and nature of using copyrighted works to train large language models (LLMs) with the goal of generating new text were deemed “transformative.”

“Like any reader who aspires to become a writer, Anthropic’s LLMs trained on such works not to replicate or replace them, but to go down a different path and create something new,” says the Court, and if this process involves the creation of copies, those copies are in turn used in a transformative and lawful way.

Each work selected for training was processed at least four times: first by copying it in digital format into the central library, then cleaning it by removing any unnecessary text, then transforming it into numeric sequences or tokens, and finally by storing the compressed copies in a dedicated repository.

No external disclosure was made, so users would never have been able to access the full texts, just as in the Google Books case. The authors do not claim any copyright infringement regarding the outputs, and all disputes remain limited to the inputs.

Regarding the books Anthropic purchased in paper format, the Court completely disagrees with the authors, stating that the simple conversion of a book into digital format is lawful and does not represent a derivative work, as in itself it has no originality, and just as a human learns from books, so too can an LLM trained on those texts.

The Pirate Copies

However, the crucial point concerns the texts that were taken from the Internet or from databases without buying printed copies or requesting the relevant rights.

On this matter, the Court’s decision was equally firm and contrary to Anthropic, which tried to defend itself by arguing that acquiring pirated copies was necessary to facilitate training or because the original texts were difficult to find.

But the Court reiterated that no ruling exists that supports the idea that “pirating” a book is reasonably necessary to write a review, conduct research, or train an LLM. This behavior would be unlawful even if the pirated copies were used immediately for transformative purposes and then deleted, but in this case the situation is worsened by the fact that Anthropic did not just use those copies for training but kept them in its central library.

Pirating works to build a research library without paying for them and keeping them in case they may prove useful was not considered a transformative use.

The same copy can be used in different ways, with different legal outcomes, and in this case Anthropic played a bad hand by declaring that the copies were obtained to “build a research library” containing well-organized facts, analyses, and expressive examples for different uses, one of which was training.

The fact that the main goal was to create this library led the judge to not consider fair use applicable for the works acquired illegally, concluding that subsequent payment cannot undo the damage caused by initial piracy.

Anyone who buys a book no longer has to justify possessing the copy, but anyone who copies the book from a pirate site has already infringed copyright.

 

Laura Turini