BOLDERROR

Artificial Intelligence Daily edition

ARTIFICIAL INTELLIGENCE AI GOVERNANCE

Washington backs OpenAI, but the case is not closed

The US government has asked the court to deem the use of copyrighted works for training large language models as fair use. This is a significant political endorsement, but it is neither a legal victory nor a universal licence to copy without limits.

By Rubén Campoy3 min read
Two professionals reviewing newspapers and archive documents on a library table.
Exclusive editorial image · BOLDERROR

The US government has formally entered the dispute between The New York Times and other publishers, and OpenAI and Microsoft. It has done so via a brief filed with the federal court in Manhattan, arguing that training large language models is, broadly speaking, a highly transformative use of original works. The administration adds an argument that shifts the tone of the discussion: restricting this practice, it claims, could harm US scientific progress, economic prosperity and national security.

It is important to clarify from the outset what has and has not occurred. The Department of Justice has not issued a ruling, nor can it. Its brief represents the government's position, which the judge may consider, but the case remains open and the decision rests with the court. Nor does it resolve the entire data supply chain in one fell swoop. Obtaining a copy, storing it, using it during training, and generating output that commercially substitutes the original are all distinct actions. A solid defence at one stage does not automatically negate the risks at others.

A model’s market position determines fair use

The legal crux is fair use, the US doctrine that balances factors such as transformative purpose, the nature of the work, the amount used, and market harm. OpenAI maintains that a model does not store a conventional archive, but rather learns statistical relationships to produce new results. Publishers counter that their articles were used without permission to create products that can answer the same questions and compete for attention, subscriptions and advertising. This is the real point of friction: not only what the training process technically does, but which market the resulting system ultimately occupies.

The federal intervention is significant because it turns an intellectual property dispute into a matter of industrial policy. If training requires a licence for every work, the cost and complexity could favour the very companies with the most capital and the best data deals. If, on the other hand, a very broad exception is granted, creators and media outlets may lose their bargaining power over the material that feeds commercial products. There is no clean solution: a poorly calibrated barrier could stifle competition, and a wide-open door could erode the production of the reliable sources that models need.

For a company deploying AI, the sensible approach is not to assume a free-for-all. It must maintain an inventory of its data sources, separate public content from material obtained under contract, document the purpose of each corpus, and control outputs that might reproduce lengthy excerpts. It is also wise to distinguish between training a model, fine-tuning an existing one, and building an application with retrieval-augmented generation. A RAG system that delivers internal documents almost verbatim has a different risk profile from a model trained to generalise from patterns.

Even greater caution is required in Spain and the European Union. The US doctrine of fair use does not automatically translate to the European framework, which operates with specific exceptions, reservations of rights, and transparency obligations. A multinational organisation needs a policy for each jurisdiction and contract, not a single global slide deck declaring the problem solved. Moreover, the publishers' lawsuit is not unique: authors, music labels and other rightsholders are pushing for different interpretations, and the first courts to address model training have not drawn a consistent line.

The key things to watch now are the judge's ruling on the merits of the case, the distinction drawn between training and market substitution, and the emergence of any licensing models that prevent data access from becoming the exclusive preserve of a few tech giants. Until then, the practical course of action is rather traditional: know what has been copied, why, with what permission, and what might come out at the other end. The novelty is the scale; the responsibility for documentation is unchanged.

End of article

Tags

  • OpenAI
  • Microsoft
  • The New York Times
  • Copyright
  • Fair use
  • AI governance

BOLDERROR Daily edition Rubén Campoy

Related

Back to the front page