Two new publishers in a battle already shaping generative AI
The Seattle Times and Newsday have in turn taken legal action against OpenAI and Microsoft, TechCrunch reports. The two publications accuse the technology groups of using their journalistic content to train artificial intelligence systems. The proceedings are part of a legal sequence that has now become central both to the press industry and to generative AI: who can use texts produced by newsrooms, under what conditions and for what compensation?
The announcement is not merely another episode in a series of similar lawsuits. Above all, it confirms a gradual shift in the balance of power between platforms and media outlets. For years, a considerable portion of the web was treated as a vast resource available for indexing, archiving, analysis or algorithmic recommendation. The rise of large language models has made this model far more questionable in publishers' eyes: their archives no longer merely help direct a reader to an article; they can help directly produce a written answer, sometimes without referring back to the original source.
In the case mentioned by TechCrunch, Seattle Times and Newsday are targeting OpenAI, the creator of ChatGPT, and Microsoft, OpenAI's strategic partner and a major player in the commercialization of generative AI services. Both companies already face several lawsuits from rights holders, including authors, artists, developers, newspapers and other content producers. The proceedings differ in their arguments and jurisdictions, but they rest on a shared question: does training a model on protected works constitute a use requiring authorization?
News publishers are in a particular position. They do not claim rights only in isolated texts: every day, they produce structured, verified, dated and ranked corpora associated with headlines, photographs, metadata and an editorial reputation. These corpora have their own economic value, especially because they are regularly updated and cover local, national and international events. For a conversational AI, this material can be particularly useful, both in historical data and in news content.
The Seattle Times and Newsday lawsuit therefore comes at a time when access to journalistic content can no longer be regarded by publishers as a mere technical consequence of being on the web. The debate goes beyond downloading or copying articles. It also concerns the market effects of products capable of synthesizing information, answering requests expressed in natural language and, potentially, keeping users within an interface rather than directing them to the site that funded the journalistic work.
The legal characterization of these practices nevertheless remains open. The lawsuits brought do not mean that courts have already ruled in favor of publishers. They do reveal, however, that holders of editorial catalogs no longer want to let this issue be decided solely by the technical choices of model providers. For AI companies, the stakes are just as significant: a general obligation to negotiate access to content, or judicial restrictions on certain training practices, could alter the costs, timelines and development strategies of their systems.
From web indexing to model training: what has changed
The current conflict is part of a long history of sometimes ambivalent relationships between the press and digital platforms. For years, search engines sent traffic to news sites while organizing access to their content through results, snippets and links. Social networks subsequently took on a decisive role in the distribution of information, with major effects on audiences, subscriptions and advertising. Generative AI adds a change in nature: rather than ranking or relaying information, it can rephrase it and present it as a standalone result.
This distinction is essential to publishers' arguments. A link generally leads to a page that retains its context, advertising, potential subscription and editorial identity. An answer generated in an AI interface, on the other hand, can satisfy a request without generating a visit. It can also reuse elements, wording or summaries from publications without the user always clearly distinguishing the original work, its date, its author or the limitations of the available information.
Language models do not work like conventional databases: they learn patterns from very large sets of texts. But this technical difference does not close the copyright debate. Opponents of these practices argue that models rely on copies made during the training phase and that they may, in certain circumstances, reproduce protected content or generate very similar versions of it. AI companies, for their part, have argued in several cases that training constitutes transformative use and should not be treated as mere editorial exploitation of every work included in the data.
U.S. courts will have to examine these arguments based on the facts specific to each case. This is one reason why press lawsuits are being closely watched. They lie at the intersection of several areas: copyright, competition, how models work, access to online content and the economic effects of new digital products. A decision concerning copying during training, the possible memorization of passages or a chatbot's ability to replace reading an article could have consequences far beyond the two titles cited by TechCrunch.
The New York Times case helped make this conflict especially visible. In December 2023, the newspaper sued OpenAI and Microsoft, alleging, among other things, infringement of its rights in its content. That lawsuit gave strong symbolic weight to the media challenge: it came from an organization with extensive archives, a powerful brand and an established digital strategy. Since then, actions brought by other news producers have reinforced the idea that this is not an isolated dispute between a technology company and a major U.S. newspaper.
The Center for Investigative Reporting, which notably publishes Mother Jones, also initiated proceedings against OpenAI in 2024. Other titles and press groups have pursued litigation, while some have favored commercial agreements. This coexistence of strategies is important: it shows that publishers do not form a homogeneous bloc. Major groups with extensive archives, negotiating capacity and international brands are not in the same position as regional newspapers, investigative media, local newsrooms or independent publications.
Seattle Times and Newsday give this discussion an additional dimension precisely because of this. Local and regional press outlets depend heavily on their ability to monetize information that often exists nowhere else: municipal affairs, schools, transportation, local justice, sports, local economy, real estate or local elections. When this data is reused, synthesized or absorbed by technological intermediaries, the issue concerns not only the abstract right to use a text. It concerns the very funding of costly journalistic coverage that is difficult to replace.
It is also necessary to distinguish model training from access to recent information. A company may train a model on a historical corpus, connect a tool to up-to-date sources or incorporate search results into an answer. These mechanisms are not identical, and the rights that may be invoked can vary depending on the uses. For publishers, this granularity does not remove the main issue: every stage of the chain, from collection to the final result, can help create value from editorial work.
An industry wavering between lawsuits and commercial licenses
The landscape is not limited to lawsuits. OpenAI has entered into several agreements with press and media groups, illustrating another possible path: licensing negotiations. The Associated Press announced an agreement with OpenAI in 2023 allowing the AI group to access part of its archives. Axel Springer subsequently announced a partnership with OpenAI concerning the use of content from its publications in the U.S. company's products. These agreements showed that news content could become an explicit commercial asset in the model economy.
In 2024, OpenAI also announced partnerships with the Financial Times, Prisa Media, which notably publishes El País, and Le Monde. These announcements were scrutinized far beyond the countries concerned. They suggest that part of the sector is seeking to turn a relationship historically based on open access to the web into contracts covering the use of archives, the display of content or the presence of quotations and attributions in AI services.
The financial terms of these agreements are generally not made public. This opacity limits comparisons and fuels a recurring concern among smaller publishers: if licenses become the primary compensation mechanism, who will be able to negotiate on fair terms? A large international group can offer thousands of publications, extensive documentary collections, legal teams and an established audience. A local media outlet may have rare and highly useful information, but lack the resources to negotiate, monitor uses and enforce its rights.
Disputes such as the one involving Seattle Times and Newsday therefore indirectly strengthen the negotiating value of the press as a whole. A lawsuit does not automatically create a license, and a commercial agreement does not amount to recognition of a general legal obligation. Nevertheless, the multiplication of proceedings increases regulatory and legal risk for AI companies. It also encourages them to clarify their data sources and demonstrate that they have sustainable access to high-quality content.
For OpenAI, these cases are unfolding as its products have become a public benchmark for generative AI. ChatGPT brought conversational interfaces into consumer and professional use, while Microsoft has integrated AI features into several of its products and services. This visibility makes both groups particularly exposed to criticism over training data. It also explains why proceedings targeting these companies are watched as tests for the entire sector, including players developing competing models.
Microsoft occupies a specific place in this matter. Its partnership with OpenAI, its cloud infrastructure and its AI-integrated products place it at several levels of the value chain. For publishers, responsibility is therefore not necessarily limited to a model's designer. It may concern companies that fund, host, distribute or incorporate these technologies into tools used on a large scale. Determining where each participant's responsibility ends will be one of the legal and economic issues to watch.
This situation contrasts with the former model of news aggregation. The press had already secured legislative changes and agreements in several countries in an attempt to rebalance its relationships with major platforms. Generative AI does not replace those debates, but extends them. It introduces a deeper question: is value created when content is published, when it is distributed to a reader, when it is used to improve a model, or when it contributes to an answer generated years later?
Licenses can provide a pragmatic answer, but they do not resolve all difficulties. They must define the content concerned, the duration of access, permitted uses, attribution, withdrawal arrangements, retention of copies and the issue of data already ingested. They may also raise the issue of updates: a press archive may be corrected, supplemented or removed for editorial or legal reasons. In an environment of models trained at scale, the ability to comply with these changes becomes as much a technical issue as a contractual one.
Copyright in the face of models: uncertainty extending beyond the United States
U.S. proceedings are particularly visible because OpenAI and Microsoft are based there and because U.S. fair use law holds a major place in the debates. But the issue is global. Data used by models is often collected from sites accessible from many countries, users query tools in several languages, and answers may be distributed in jurisdictions with different rules. For European and French-speaking publishers, U.S. cases are therefore important indicators, without by themselves determining the legal answer applicable in their territory.
In the European Union, the directive on copyright in the Digital Single Market introduced a framework concerning text and data mining. Its regime distinguishes in particular between mining conducted for scientific research purposes by certain organizations and institutions, and other text and data mining uses. For the latter, rights holders can reserve their rights under certain conditions. In the case of content made available online, this reservation must in particular be able to be expressed in an appropriate manner, including through machine-readable means.
This European framework is often presented as a possible tool for publishers wishing to refuse certain automated uses of their content. But its practical implementation is far from simple. A reservation expressed on a website, in terms of use or through technical instructions does not automatically answer all questions: how can it be proven that it was taken into account? From what date does it take effect? How should copies already collected be handled? And how can systems whose datasets are not always published with enough detail for independent verification be monitored?
In France, these issues intersect with a press sector already familiar with negotiations with platforms. The creation of a neighboring right for press publishers and agencies, following on from the European directive, has placed compensation for the reuse of news content at the heart of relations between media outlets and digital players. The neighboring right is not the same as the debate on model training. It nevertheless attests to an economic principle that has become decisive: producing professional information does not mean that information can be used without consideration by any player with distribution or computing infrastructure.
For French publishers, the U.S. precedent may have a market effect even without direct legal force. If AI companies begin multiplying licenses with U.S., British or European groups, French-language titles will be better able to argue that their corpora should likewise be integrated into a contractual framework. Conversely, if courts recognize broad freedom to train on publicly accessible works, negotiating opportunities could be narrower, especially for publishers without an international brand.
Language is an important factor. French-language content is not a merely interchangeable subset of the English-language web. It provides knowledge related to French, Belgian, Swiss, Quebec, African and European institutions, as well as to specific political, legal and cultural contexts. For an assistant intended to respond accurately to French-speaking users, access to quality French-language sources can improve the accuracy, vocabulary, contextualization and currency of answers. This specific value strengthens the argument of publishers who refuse to see their archives regarded as free raw material.
It also raises a difficulty: not all French-language newsrooms have the capacity to verify how their content circulates in global datasets. Technical blocking mechanisms, such as robots files or metadata, may signal a preference or reservation, but they rely largely on voluntary compliance by collectors. Legal action then remains a costly and lengthy last-resort instrument that not all media outlets can mobilize with the same intensity.
The debate consequently concerns public authorities, competition authorities, collective management organizations, publishers' associations and digital law researchers. It is not only about protecting existing revenues. It is also about determining whether AI development can sustainably rely on editorial production whose cost is borne by subscriptions, advertising revenue, public support or investors, while another industry derives technological and commercial value from that production.
The economic risk: can an AI answer replace a reader's visit?
One of the most sensitive points for media outlets lies in the substitution effect. When a user asks an assistant to summarize the news, explain a political decision, find a local fact or compare several events, they can obtain an answer without opening the publisher's site. This experience is useful for the user, but it may reduce opportunities for direct visits. Yet visits often remain tied to advertising, conversion to a subscription, discovery of other articles and brand exposure.
The consequences are not identical for all content. A general query about an old subject does not necessarily replace reading an investigation, report or in-depth analysis. But the accumulation of practical or factual requests can divert part of an audience that previously came through a search engine. The problem becomes more acute if the conversational tool is perceived as a final destination rather than a gateway to identified sources.
Publishers do not necessarily challenge the idea that a technological tool can guide, summarize or help users discover the news. Licensing agreements even show that cooperation is possible. The point of friction lies in control. A newsroom must be able to choose whether its content is used, know the terms of that use and obtain compensation when that use feeds a commercial product. From this perspective, the Seattle Times and Newsday dispute also concerns a principle of data governance: should the information producer have a genuine right to negotiate the place of its work in AI systems?
The quality of information is another economic issue. Generative models can produce fluid answers, but they can also be wrong, lack context or present outdated information. For media outlets, the use of their work by these systems therefore raises a particular tension: their content may help improve the apparent credibility of an answer while being dissociated from the editorial mechanisms that normally ensure its verification and updating.
Attribution appears to be a partial answer. When the tool cites a publication and links to its article, it makes the origin of the information more visible and may generate traffic. But attribution does not automatically resolve the issue of copies used for training, nor that of compensation. Nor does it guarantee that the reader will follow the link. Publishers will therefore have to balance several sometimes conflicting objectives: visibility, control, direct revenue and preservation of the relationship with their audience.
The litigation also has implications for the corporate customers of AI providers. Many organizations already use generative tools for monitoring, document summarization, customer service, writing or analysis. If rules governing access to content become stricter, these users will need to pay closer attention to the provenance of data, contractual guarantees and the risks associated with using results produced by models. The issue no longer concerns only the labs that train systems, but all the players that deploy them in their processes.
For media outlets, the challenge is ultimately strategic. Refusing all access may reduce their presence in new information interfaces. Accepting access without conditions may weaken their economic model and their control over archives. Between these two positions, licenses, APIs, display partnerships, citation systems and opt-out mechanisms outline a range of possible solutions. The Seattle Times and Newsday lawsuit nevertheless reminds us that not all of these solutions will advance through consensus: the law is also used to impose negotiation where commercial power imbalances exist.
Toward a more closed market for editorial data
The most important prospect is that of a lasting transformation of the data market. Large models have long been associated with the idea of mass web collection. The lawsuits filed by publishers, including the one reported by TechCrunch involving Seattle Times and Newsday, instead push toward an environment in which the origin, licensing and traceability of content become much more visible commercial and legal criteria.
Such a movement could favor data providers with structured archives, metadata expertise and clearly identified rights. News agencies, major publishing groups, specialized databases and media outlets with historical collections could negotiate a more important place in the AI value chain. But there is a corresponding risk: a market based on large licenses could concentrate revenue among the most powerful players, while smaller newsrooms would have greater difficulty having the value of their publications recognized.
For the French-speaking market, the challenge will be not to let linguistic and editorial quality become a free externality of technological development. French and European newsrooms have content necessary for an AI capable of understanding public debates, regulations, institutions and local realities. The challenge will be to organize this value without locking down innovation or depriving citizens of new tools for accessing information.
U.S. court decisions will not alone determine this evolution. European legislators, regulators, licensing agreements, companies' technical choices and user practices will also matter. But the accumulation of lawsuits is already changing the framework of the debate. It compels AI players to recognize that journalistic content is not merely available data: it is protected work, economic assets and the result of an editorial activity whose sustainability also determines the quality of the information on which future models will seek to draw.
Comments· 1 comment
Thank you for covering this important development. It’s encouraging to see the debate around journalism, licensing, and AI being taken seriously.