Legal & Compliance | 4 min read

US Appeals Court Upholds: Training AI on Copyrighted Material Is Not Fair Use — Ross Intelligence Loses

A US appeals court ruled that training an AI model on copyrighted content is not fair use, creating binding precedent that threatens the training data strategies of major AI developers.

Hector Herrera
Hector Herrera
A law office where a person is reviewing related to US Appeals Court Upholds: Training AI on Copyrighted Materia
Why this matters A US appeals court ruled that training an AI model on copyrighted content is not fair use, creating binding precedent that threatens the training data strategies of major AI developers.

A US appeals court upheld a lower court ruling this week that AI legal research firm Ross Intelligence's use of copyrighted legal content to train its AI model did not constitute fair use — creating binding appellate precedent against the training data strategy that underlies most of the modern AI industry.

The ruling is not advisory. In the circuits covered by this court, AI companies that trained models on scraped licensed content now face renewed and legally strengthened exposure. Every foundation model developer with unresolved questions about training data provenance should be reviewing their position with legal counsel.

What the Case Was About

Ross Intelligence built an AI legal research tool designed to compete with Westlaw, Thomson Reuters' flagship legal research platform. To train its model, Ross used Westlaw headnotes — short editorial summaries that distill the legal significance of court opinions into a searchable format.

These headnotes are not the underlying court opinions, which are public domain. They are Thomson Reuters' own copyrighted editorial work, produced by lawyers who review each opinion and craft language identifying its precedential value. The copyright isn't in the law — it's in the editorial layer applied to the law.

Thomson Reuters sued for infringement. Ross argued that training an AI model on copyrighted content to build a competing product constituted fair use — a legal doctrine that permits copying under specific conditions without the rights holder's permission.

Why the Court Said No

The lower court found, and the appeals court agreed according to Futurism's reporting, that Ross failed to satisfy the key factors courts evaluate for fair use:

Purpose and character of the use: Ross's use was commercial and directly competitive with the original product. Fair use favors transformative uses that add new meaning or value; using copyrighted material to train a product designed to replace the original does not qualify as transformative in the legally meaningful sense.

Effect on the market: The use substituted for a licensing market Thomson Reuters had established. Westlaw headnotes are commercially licensed. Ross's AI tool didn't reference or critique the headnotes — it used them to displace demand for the licensed product. Courts weigh this heavily.

Nature of the copied work: The headnotes are creative editorial expression, not raw factual data. Factual compilations receive less copyright protection; expressive editorial work receives full protection.

The combination of commercial purpose, competitive substitution, and copying of expressive content — rather than unadorned facts — made the fair use defense legally untenable.

What This Means for Foundation Model Developers

The Ross ruling creates binding precedent in the Third Circuit, which covers Pennsylvania, New Jersey, and Delaware — jurisdictions that include major technology, financial, and legal industry activity. It will be cited as authority in dozens of pending AI copyright cases, including suits against OpenAI, Anthropic, Google, Meta, and Stability AI.

The critical distinction the court drew is important for understanding who faces the greatest exposure: the problem isn't simply that copyrighted material was used for training — it's that the AI output substituted for a licensing market the copyright holder had established. Companies that trained on datasets where the rights holders had commercially licensed the content for similar purposes face the most direct application of this ruling.

This includes:

  • Legal research content (the direct subject of this case)
  • Licensed news archives
  • Academic and scientific publishing datasets
  • Financial data with commercial licensing programs
  • Any professional content category where a licensing market for AI training has been established or could reasonably have been established

Companies whose training data came primarily from openly licensed sources, public domain materials, or legitimately licensed datasets have more defensible positions. But demonstrating that provenance — in discovery, in litigation, to courts — is now a legal priority, not just an operational one.

The Training Data Market Accelerates

One effect of this ruling is already visible in the broader market: AI companies have been signing training data licensing deals at an accelerating pace throughout 2026. OpenAI has agreements with Associated Press, Axel Springer, and multiple book publishers. Google has established licensing relationships with major media organizations. The rate of new training data licensing agreements will likely accelerate further as the legal risk of unlicensed commercial content in training datasets crystallizes into binding precedent.

For larger AI companies, licensing at scale is expensive but achievable. For smaller AI companies and startups building on open-weight models, the economics are harder. The cost of licensing the breadth of training data required to build a competitive general-purpose model is significant. This ruling doesn't make it impossible to train AI models on commercial content — but it establishes that doing so without a license carries real legal risk, not just theoretical exposure.

The practical pressure will favor companies that built training data acquisition strategies around licensing agreements, synthetic data generation, or public domain sources — and disadvantage those that relied on broad web scraping of commercially licensed content.

What to Watch

Whether Ross Intelligence or other AI defendants in pending copyright cases petition the Supreme Court for certiorari (review) is the critical next step. If the Supreme Court declines to take up the issue, the Third Circuit precedent stands and begins to function as a de facto national standard — courts in other circuits frequently follow well-reasoned appellate decisions even when not bound by them. A cert petition signals that the legal community views the stakes as large enough to warrant definitive national resolution. Given the scale of pending AI copyright litigation, a high court decision in this area within the next two to three years is plausible.

Key Takeaways

  • ✓ Purpose and character of the use
  • ✓ Effect on the market
  • ✓ Nature of the copied work

Did this help you understand AI better?

Your feedback helps us write more useful content.

Hector Herrera

Written by

Hector Herrera

Hector Herrera is an AI systems architect in Houston and founder of Hex AI Systems. He designs and runs AI systems in production and writes daily about how AI is reshaping business, government and everyday life. 20+ years building for the web. Houston, TX.

More from Hector →

Get tomorrow's AI briefing

Join readers who start their day with NexChron. Free, daily, no spam.

More from NexChron