Washington backs OpenAI against the New York Times
The lawsuit pitting the New York Times against OpenAI and Microsoft is no longer being fought solely between a major press group, a software publisher, and the creator of ChatGPT. According to a report by TechCrunch, the U.S. government has intervened in favor of OpenAI on an issue that now concentrates a large share of the legal tensions surrounding generative artificial intelligence: under what conditions can language models be trained on works protected by copyright?
The political scope of this intervention goes beyond the New York newspaper's case. By defending an interpretation favorable to training large language models on publicly accessible content, the federal executive is advancing a strategic argument: the ability of American companies to remain competitive in AI should not be hampered by an interpretation of copyright that could limit access to the data corpora needed for research and development.
This position does not settle the dispute. The decision still rests with the judge hearing the case. Nor does it erase the claims of the New York Times, which accuses OpenAI and Microsoft of using its content without authorization and argues that their tools may compete with certain uses of its journalism. But Washington's intervention brings an unusual dimension to the debate: the federal government is no longer merely observing a private dispute over the use of protected texts. It sees it as a matter of industrial policy, global technological competition and, implicitly, U.S. digital sovereignty.
For publishers, authors, artists, photographers and other rights holders, the signal is significant. Since the mass rollout of generative tools beginning in 2022, litigation has multiplied in the United States. Plaintiffs are asking courts to determine whether absorbing immense quantities of text, images, code or sound into training datasets constitutes infringement, transformative use covered by fair use, or a practice that should be subject to licensing and compensation.
For AI companies, by contrast, the issue is existential. Modern generative models are built on considerable volumes of data. A general requirement to negotiate individual authorization in advance with every rights holder would profoundly change the sector's economics. It would also raise practical problems: identifying all rights attached to every document, knowing the exact provenance of data already incorporated into historical corpora, and distinguishing content that can be freely reused from content that cannot.
The New York Times lawsuit is therefore becoming a major political test. It is not only a question of whether a media outlet's articles were used lawfully. It is also about determining whether U.S. copyright law, shaped long before the arrival of large language models, can protect creators without preventing domestic companies from developing a technology Washington considers strategic.
A landmark dispute between the New York Times, OpenAI and Microsoft
The New York Times brought its action against OpenAI and Microsoft at the end of 2023. The newspaper argues in particular that the defendants reproduced and exploited its journalistic works without authorization in order to develop and market products based on artificial intelligence. Microsoft is targeted because of its close partnership with OpenAI and its integration of the latter's technologies into several of its services.
The complaint immediately attracted attention, both because of the plaintiff's stature and the nature of the allegations. The New York Times is not merely a content producer: it is a global media institution with substantial archives and a digital business based on subscriptions. Its economic interest is directly tied to the value of its texts, investigations, reviews and editorial work. Its action thus gave a concrete face to the conflict between the press and the labs that train models on the Web.
OpenAI has challenged the allegations against it. The company has argued in particular that some examples cited by the newspaper to show reproduction of its articles resulted from prompts specifically designed to push the models to reproduce passages in an abnormal manner. The dispute therefore concerns both training data and the outputs produced by the systems: can a response resembling an article demonstrate unlawful reproduction, and at what frequency or level of fidelity?
This distinction matters. Training a model consists of adjusting its parameters using very large datasets. The result is not, in principle, a database that makes it possible to consult every source document as in a conventional archive. Rights holders nevertheless respond that this technical difference is not enough to eliminate the use of their works, nor the economic risks if a tool can produce summaries, responses or passages that replace access to the original.
The New York Times has also argued that AI products could divert part of the audience and revenue associated with access to its content. This question of substitution lies at the heart of the debates. A conversational assistant that answers a factual question by drawing on multiple sources does not necessarily have the same effect as a system capable of providing, on demand, the text or a highly detailed summary of an article reserved for subscribers. Between these two situations, courts must assess the nature of the use, its transformative character and its potential effect on the work's market.
The case is part of a broader wave of proceedings. Authors have sued OpenAI and other technology players over the use of books; visual artists have brought actions against companies specializing in image generation; Getty Images sued Stability AI; Universal Music Group, Sony Music Entertainment and Warner Records brought actions against Suno and Udio over music generation. Each case has its own facts, works and arguments, but all raise a common question: can the development of a generative system rely on protected creations without a prior contract with their rights holders?
The New York Times case nevertheless stands out because of the importance of journalism in American democracy and the particular relationship between information, subscriptions and online search. For years, press publishers have negotiated with digital platforms over visibility, indexing and compensation for their content. The emergence of AI assistants shifts this balance of power. Whereas a search engine generally directs users to a source, a chatbot can provide an answer directly in its interface. The value of the original page, the visit and the relationship with the reader then become more difficult to preserve.
OpenAI, for its part, has entered into content agreements with several press groups and media outlets, including Associated Press, Axel Springer, Le Monde, the Financial Times, News Corp and Condé Nast. These partnerships show that a contractual path exists. They do not, however, resolve the general legal question. An agreement with certain publishers does not mean that OpenAI accepts that all protected data requires a license. Conversely, entering into contracts may be interpreted by rights holders as recognition of the economic value of their catalogs.
It is in this environment that the federal intervention reported by TechCrunch takes on its full significance. Washington is not simply ruling on the value of press content. It is defending a view under which rules applicable to training must take account of the public interest associated with maintaining a powerful and innovative American AI ecosystem.
Copyright as an instrument of technological competitiveness
In the United States, much of AI companies' argument rests on fair use, a copyright exception assessed on a case-by-case basis. It does not operate as an automatic permission. Courts traditionally examine several factors, including the purpose and character of the use, the nature of the protected work, the amount used and the effect of the use on the work's potential market.
The notion of transformation is particularly important. Under U.S. case law, a use may be considered more favorable to fair use when it adds a new purpose or character rather than simply substituting for the original work. AI companies argue that training transforms data into statistical capabilities: the model learns patterns of language, style, syntax or knowledge without constituting a library of works ready to be redistributed.
Plaintiffs do not necessarily dispute that models produce new results in many cases. Their argument instead concerns the initial extraction, the considerable quantity of content used and the possible economic harm. For a publisher, the question is not solely one of verbatim textual copying. If a system can answer internet users' requests based on the substance of its work, it may reduce the incentive to consult or purchase that work. And if responses reproduce passages, even exceptionally, that may increase the risk.
According to TechCrunch, the U.S. government is highlighting a concern about competitiveness. The idea is clear: overly restrictive rules on data access could disadvantage American companies against foreign competitors. In the race for large models, computing power is crucial, but so is data. The most capable foundation models are generally trained on very large corpora made up of text, code, images, multimedia content and other resources.
This reading turns copyright into an industrial-policy variable. Historically, U.S. copyright seeks to encourage creation by granting authors time-limited control over the exploitation of their works. The U.S. Constitution itself links the protection of authors and inventors to the promotion of progress in science and useful arts. With generative AI, the two objectives may come into tension: protecting the ability of media outlets and creators to finance their production, while avoiding reserving access to knowledge and corpora for a handful of players able to pay for global licenses.
Supporters of an approach favorable to training also stress the risk of concentration. If every quality dataset had to be contractually licensed, companies that are already wealthy and established could gain an even greater advantage. Startups, university labs and nonprofit organizations would be less able to assemble comparable corpora. This argument does not imply that the use of protected content is necessarily lawful, but it fuels the debate over the economic consequences of a mandatory licensing regime.
Conversely, representatives of creators believe that a system without compensation can also reinforce concentration. Major labs have computing infrastructure, research teams and access to massive data. Without an obligation to negotiate, they can absorb the value created by media outlets, writers, photographers or musicians without sharing the revenue generated by their products. The question then becomes one of a transfer of value between cultural industries and technology platforms.
The federal government is not presenting itself here as a neutral arbiter of this redistribution. Its intervention reflects a hierarchy of priorities in which U.S. leadership in AI occupies a central place. This direction is part of an international context in which the United States, China and Europe are each seeking to influence artificial intelligence infrastructure, models, chips and standards.
For OpenAI, whose tools have become one of the symbols of consumer-facing generative AI, this political support is significant. It does not guarantee the outcome of the lawsuit and does not amount to general validation of all data collection or training practices. But it gives weight to the argument that courts should examine these practices in light of their effects on innovation, research and the United States' international position.
A legal battle reshaping relations between AI and the media
Washington's intervention comes as relations between AI labs and publishers fluctuate between litigation and negotiation. On the one hand, lawsuits seek recognition of copyright violations and compensation, or even restrictions on the use of certain works. On the other, commercial agreements seek to organize access to selected content, sometimes with attribution, citation or distribution mechanisms within AI products.
This coexistence of strategies is revealing. Publishers do not form a homogeneous bloc. Some favor partnerships, believing they can create new revenue in an environment where AI is becoming a gateway to information. Others fear that contracts will compensate for only a small part of the value extracted, or that conversational tools will weaken the direct connection with the public. The same press group may, moreover, seek to negotiate certain uses while remaining vigilant about the risks of reproduction and substitution.
The precedent of search engines is often cited, but it has its limits. Search engines profoundly changed the economics of the press by capturing part of attention and advertising while sending traffic to publishers. Conversational systems can retain links and display sources, but they can also reduce the need to click. The degree of transparency about the origin of responses, the prominence given to citations and users' ability to trace a response back to an article are therefore important elements for media outlets.
Responses produced by models also raise an issue of reliability. Journalism is not limited to an accumulation of facts. It relies on prioritizing information, a verification method, editorial context and legal accountability. When an assistant summarizes a current-affairs topic, it may omit nuances, mix up periods or produce errors. For publishers, distributing responses without a visit to their site can therefore pose an economic problem, but also a reputational problem when their information is distorted or attributed imprecisely.
Labs have an interest in reducing these frictions. Access to quality, recent and verified content is valuable for products that aim to answer information queries. The agreements OpenAI has entered into with press organizations point in this direction. They may make it possible to power certain search experiences or provide more directly identifiable sources. However, these contracts are not sufficient to resolve the debate over historical datasets that contributed to the training of models that already exist.
The New York Times case also involves a complex procedural and evidentiary question. Large models are opaque systems made up of a very high number of parameters. Establishing that a specific text was included in training, then demonstrating a connection between that inclusion and a given output, is not straightforward. Companies may invoke trade secrets; plaintiffs seek information on corpora, filtering methods and measures taken to limit the memorization or reproduction of works.
The final decision, whatever it may be, could therefore rest as much on highly technical facts as on legal principles. Courts will have to assess the characteristics of the models at issue, the observable behavior of products, the safeguards implemented and the possibility of harm to the markets for works. A clear victory for either side cannot be presumed from the government's intervention alone.
There is nevertheless an immediate effect: the federal position makes it more difficult to view the issue as a mere commercial conflict between private companies. By invoking the competitiveness of American industry, Washington asserts that case law unfavorable to generative models could have consequences extending beyond the revenue of a particular publisher. Rights holders may respond that weakening content protection would likewise have collective consequences for the production of information, culture and knowledge.
This opposition cannot be reduced to a choice between innovation and creation. AI innovation itself depends on content created by humans, often by sectors already under strong economic pressure. Creation, in turn, can benefit from new tools for research, translation, assistance or dissemination. The real challenge is to define legal and economic conditions under which these two activities can coexist without one permanently destabilizing the other.
For France and Europe, a U.S. case with direct consequences
The lawsuit is taking place in the United States and falls under U.S. law, but its repercussions are being closely watched in Europe. Large language models marketed globally are often developed by American companies. U.S. court decisions can influence their data practices, partnership strategies, content-removal mechanisms and position in negotiations with European publishers.
The European framework nevertheless differs significantly from the U.S. approach. The European Union introduced, in the Directive on Copyright in the Digital Single Market, exceptions relating to text and data mining. One concerns scientific research; another permits text and data mining under certain conditions, particularly when rights holders have not expressly reserved that use. This opt-out mechanism is often central to discussions between AI players and rights holders.
France has transposed this directive into its national law. For French publishers, producers and creators, the practical issue is how to effectively express a reservation of rights in a digital environment where content circulates, is copied, indexed and sometimes incorporated into corpora by multiple intermediaries. A theoretical reservation is less useful if it is not technically readable, legally clear and effectively respected by data collectors.
The European Union also has its regulation on artificial intelligence, the AI Act. This text provides, in particular, transparency obligations for providers of general-purpose AI models, including making available a sufficiently detailed summary of the content used for training, in accordance with the arrangements defined by the European framework. It does not create a general license for the benefit of rights holders and does not resolve all data disputes. But it seeks to make the ecosystem less opaque than the compilation of many training corpora has historically been.
The comparison with the United States reveals two legal philosophies. U.S. fair use is a flexible doctrine, assessed by judges and heavily dependent on context. European law relies more on exceptions defined by legislation, including text and data mining, as well as on the possibility of reserving certain uses. In both cases, the outcome remains uncertain for economic players: AI is evolving rapidly, while court decisions and implementing rules move at a slower pace.
For French media outlets, the issue is all the more important because the online information market already depends heavily on platforms, search engines and social networks. The arrival of conversational interfaces adds a potential intermediary between the newspaper and its reader. If these interfaces answer without directing users to sources, publishers risk losing part of the traffic and associated revenue. If they clearly cite sources and establish distribution agreements, they may also become a complementary channel, provided the distribution of value is considered acceptable.
French-speaking press groups also have a particular interest in the linguistic quality of models. Dominant models have long been heavily structured by English-language data. Content in French, regional languages or other European languages may be underrepresented in certain corpora, even though it is essential for developing effective assistants for local audiences. An overly costly licensing policy could limit smaller labs' access to these resources; a complete absence of compensation could weaken the content producers on whom this linguistic quality depends.
The prospect of European technological sovereignty further complicates the equation. European companies, including several that develop language models, must navigate between compliance with European law, competition from players with considerable resources and the expectations of rights holders. For them, legal clarity on text and data mining is an issue of competitiveness. But that clarity must also allow European creators to understand when, how and under what conditions their works are used.
In this context, the U.S. intervention in favor of OpenAI should not be read as a rule applicable in France. Rather, it is a political indicator. It shows that the United States is prepared to view access to data as an element of its power in AI. Europe will have to determine how far it wishes to prioritize the competitiveness of its own players without abandoning the protection and compensation mechanisms on which its cultural industries rely.
The lawsuit as a laboratory for a future economy of cultural data
Washington's position may influence arguments made in other proceedings, but it will not eliminate the need to examine each case according to its facts. Disputes relating to books, images, music, computer code or the press do not raise exactly the same questions. The behavior of an image generator, a programming assistant and a news chatbot cannot be assessed using a single criterion.
The market could evolve toward a combination of mechanisms. Voluntary agreements between labs and rights holders may multiply for content with high commercial value, particularly archives, recent data, specialized catalogs and verified information. At the same time, companies will continue to defend the idea that part of training falls under uses permitted by existing exceptions or by fair use. Courts will determine the boundaries between these two areas.
A judicial clarification favorable to OpenAI would strengthen the position of American labs in negotiations. They could argue that a license is not systematically required for training, even if it remains desirable for access to premium or up-to-date content. A clarification favorable to the New York Times, by contrast, could increase pressure for contracts, compensation mechanisms and technical traceability solutions. In both cases, the effects would not be identical depending on the size of the players: a major technology group and a startup do not have the same capacity to absorb legal costs or enter into global agreements.
The issue of transparency will remain central. Without sufficiently precise information on collection practices, it is difficult for rights holders to assess the use of their works. Without predictable rules, it is difficult for developers to know which corpora they can draw upon. European debates over summaries of training data and U.S. legal actions show that the opacity inherited from the early years of large models is becoming increasingly untenable as these systems are integrated into mass-market commercial products.
It is also necessary to distinguish model training from model deployment. A model may have been developed from a disputed corpus while being equipped with safeguards intended to reduce the reproduction of protected content. Conversely, a model trained with licensed data may produce problematic responses if its interface encourages the reproduction of texts too close to the originals. Regulators, judges and companies will probably have to consider the entire chain: data acquisition, training, adjustment, product operation, source attribution and complaint handling.
For the press, the long-term challenge is less about defending an abstract principle than preserving the material conditions for producing information. Journalistic content is expensive to produce, especially when it relies on investigation, local expertise or coverage of specialized subjects. If AI tools capture most of the value of information without contributing to funding its production, the risk concerns the diversity of available sources. If the rules become too burdensome or too fragmented, they may also limit the emergence of innovative services useful to citizens and newsrooms.
Federal support for OpenAI thus reveals a broader reality: copyright is now being discussed as AI infrastructure. It is no longer only about protecting a work after publication, but about arbitrating access to the informational raw material that feeds generative systems. The United States appears determined to prevent a restrictive interpretation from reducing its industrial lead. Europe, for its part, seeks to reconcile innovation, transparency and creators' rights within a more codified framework.
The forthcoming judgment in the New York Times case will undoubtedly not, on its own, establish all the rules of the generative economy. But it will help define the balance of power between platforms, labs, publishers and rights holders. For French-speaking players, the challenge will be not to be subject to a standard shaped solely by U.S. courts and contracts entered into by the largest groups. The next stage will not be merely legal: it will concern the ability of media outlets, creators and European companies to negotiate data-access conditions compatible with AI that is effective, traceable and sustainably funded.
Comments· 3 comments
What specific legal argument is the administration supporting here: that training on copyrighted works is fair use, or something narrower about how the training data was obtained? I’m curious whether its position addresses the Times’ claims directly or mainly sets out a broader policy view.
My understanding from the summary is that the backing appears to signal support for a right to train models on protected material, but the exact scope would depend on the administration’s filing. It would be worth checking whether it discusses fair use explicitly and whether it distinguishes training from reproducing outputs.
The article’s framing suggests the intervention could matter beyond this one dispute, yet it does not by itself tell us which legal limits the administration accepts. The key details would likely be whether the position addresses copying during training, market harm, and safeguards against models reproducing protected text.