“A smart man knows when he is right; a wise man knows when he is wrong.”
With these words, Judge Bibas of the Delaware District Court opens his latest Opinion, in which he changes his previous stance on fair use in the case Thomson v. Ross (Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., Case No. 1:20-cv-00613, D. Del. 2021).
The Case
On May 6, 2020, Thomson sued ROSS, accusing it of copying the Westlaw legal database—including annotations—to train its own legal research system. Specifically, ROSS had used LegalEase, a Westlaw partner, through which it was able to access Thomson’s database, even though the user license explicitly prohibited its use for developing competing products—something ROSS ultimately did.
Thomson therefore asked the court to find ROSS liable for direct copyright infringement as well as license violation, claiming that ROSS used information from the database to create an AI-based system in direct competition with Thomson’s.
ROSS denied the allegations, arguing that any overlap was negligible, that the annotations were not subject to copyright protection, and—most notably—invoking the U.S. doctrine of fair use, which permits limited use of protected works without the copyright holder’s consent, provided certain conditions are met. These generally include consideration of four key factors:
- Transformative use of the second work—whether it has a new purpose or meaning compared to the original;
- The nature of the copyrighted work—particularly its originality;
- The amount and substantiality of the portion used;
- The effect on the market, especially if the second work competes with the original.
The Judge’s Decision
Initially, the judge held that no preliminary assessment could be made and that a full trial would be necessary to understand how the system operated.
Now, however, he states that after a deeper review of the case, he is ready to issue a ruling—changing his earlier position—and he rejects the applicability of fair use to this case.
Commentators have hailed this preliminary decision as a major victory for authors, but it is still too early to draw sweeping conclusions, for two main reasons:
First, although the case involves a database that uses artificial intelligence for research purposes, this is not a case involving generative AI. Second, the data in question were not used to train an AI system, so the most pressing and relevant issues are yet to be decided.
According to the judge, ROSS copied the headnotes written by Thomson for each case, classified them with numerical codes, and used them as a source to deliver search results—not to train an AI model. This is, therefore, a case of database plagiarism: the copied database was used as the foundation for a more advanced AI-powered system that enhances searches. But this has nothing to do with the very different and serious issue of using copyrighted training data.
The judge makes this clear in his ruling:
“It is undisputed that Ross’s AI is not a generative AI (i.e., one that autonomously creates new content). Instead, when a user enters a legal question, Ross returns relevant, pre-written judicial opinions. This process is similar to how Westlaw uses editorial notes and its numbering system to provide a list of relevant annotated cases.”
This is, in fact, a classic case of infringement, unrelated to the use of training data. That’s why the judge is now able to rule, setting aside his earlier doubts, and declares that ROSS’s use of its competitor’s database is unlawful and does not qualify for protection under the fair use doctrine.
This decision is far from settling the ongoing debate over the use of copyrighted data for training AI systems. Legal teams involved in other major lawsuits—such as those facing OpenAI or Anthropic—will have ample room to argue how their cases differ from this one.
The Crucial Point
There is, however, one part of the decision that could raise concern. After clarifying that ROSS did not use the data for AI training purposes, the judge adds:
“It does not matter whether Thomson Reuters used the data to train its own legal research tools; the impact on a potential market for AI training data is sufficient. Ross bears the burden of proof. It did not provide enough facts to show that such markets do not exist and would not be affected.”
The judge thus seems to suggest that it doesn’t matter whether the data were used to train an AI system (which, in this case, is not even generative AI). What matters is the potential market impact, including on a hypothetical market for training data—which ROSS failed to prove does not exist or would not be affected by its system.
This notion of a “potential market” for training data opens the door to the possibility that AI systems may need to pay for the use of such data—even though the judge stops short of explicitly stating this, and it’s far too early to know if this represents the court’s broader stance.
Whether this small fissure will grow into a crack in the applicability of fair use to training data—or remain just a minor scratch—will depend on upcoming rulings in much more significant and relevant cases still awaiting final judgment.
Laura Turini