Anthropic will ultimately have to pay $1.5 billion to end a massive dispute over the alleged use of copyright-protected books. As TechCrunch reported, a US federal judge gave final approval to the settlement reached between the company behind the Claude assistant and a group of authors. The deal is one of the costliest, and arguably most consequential, episodes in the legal battle over generative artificial intelligence training data.

The figure is striking in its scale, but it should not obscure the key limitation of this decision: the settlement specifically avoids a ruling on the merits of the issue most disputed by the industry. US courts have not, on this occasion, established a general rule stating under what circumstances training an AI model on protected works does, or does not, fall under fair use, the US copyright exception based in particular on the transformative nature of a use.

What has been approved is therefore a financial compromise, not a complete legal doctrine applicable to all generative AI. But this nuance does not diminish the significance of the case. By agreeing to pay $1.5 billion, Anthropic is sending a significant signal to the market: the provenance of datasets is no longer merely a matter of reputation, compliance or commercial negotiations with publishers. It can become a financial risk capable of directly weighing on strategy, financing, audits and the ability to launch new models.

A landmark settlement in the training data war

The case is part of a wave of lawsuits filed in the United States by authors, visual artists, press groups, music labels and rights holders against companies developing generative AI systems. Since ChatGPT’s public breakout at the end of 2022 and the acceleration of competing models, training has emerged as one of the sector’s main legal fronts: what data was used, how was it obtained, what rights cover its exploitation, and to what extent can a model trained on works be regarded as a new use rather than an unlawful reproduction?

Anthropic, founded in 2021 by former OpenAI members, is among the players that have become central to this competition. Its Claude assistant is presented as a competitor to products from OpenAI, Google, Microsoft, Meta and xAI. The company has also stood out for its messaging on system safety, notably around its so-called “constitutional AI” approach. But its development depends, like that of its rivals, on very large quantities of text and other content used to train language models.

The lawsuit targeted Anthropic over the alleged use of pirated works in its training operations. The plaintiffs argued that protected books had been obtained through pirate digital libraries, including Library Genesis, often called LibGen, and Pirate Library Mirror. These collections of files have long circulated on the internet and contain a very large number of books accessible without the authorization of their authors or publishers.

The dispute took on a particular dimension because it did not concern simply the automated analysis of lawfully obtained books. The provenance of the copies lay at the heart of the proceedings. For the authors, building or maintaining a library from pirated files could not be offset by the subsequent technological purpose. For Anthropic, the issue of fair use had to be assessed in light of the training itself and the alleged transformative nature of that operation.

Before the settlement, federal judge William Alsup had issued an important ruling in the case. He found that training models on lawfully acquired books could fall under fair use, regarding that use as highly transformative. This part of the decision was closely watched by the industry because it offered a strong argument to companies claiming that a model’s statistical learning does not amount to a simple republication or making the original works available.

But the judge also separated that reasoning from the issue of copies obtained through piracy. Under that distinction, the potentially transformative nature of training did not automatically resolve the issue posed by acquiring and retaining thousands of books from unlawful sources. That point still had to be examined as the litigation continued. It was in this context that Anthropic and the authors reached their deal.

TechCrunch describes the settlement as a major decision, rightly so. The $1.5 billion amount vastly exceeds settlements usually associated with individual copyright disputes in the digital economy. It reflects the potential volume of works involved, but also the considerable risk the company would have faced had the case gone to trial over pirated copies and potential damages.

What exactly the US judge approved

The court’s final approval makes the settlement effective within the authors’ class action. The rights holders concerned can therefore be compensated under the terms provided by the agreement, rather than having to wait for the outcome of a lengthy, uncertain trial that could potentially be followed by appeals. For Anthropic, the deal ends this specific litigation and reduces legal uncertainty whose cost could have become difficult to anticipate.

The settlement was presented as covering a considerable volume of books. Information made public during the proceedings indicated that around half a million works could fall within its scope. The $1.5 billion sum therefore corresponds to compensation that can reach around $3,000 per work concerned, subject to the precise allocation and eligibility rules set out in the agreement.

This mechanism is worth noting. In copyright litigation, the issue is not limited to immediately measurable economic harm, such as lost sales. In the United States, statutory damages can be high when infringement is deemed willful. AI companies therefore have a major interest in preventing a court from calculating, work by work, the consequences of resorting to unlawful copies. The settlement allows Anthropic to set an overall amount, even if it is exceptional.

Approval does not mean, however, that the judge concluded all allegations had been proven after a trial. A settlement is, by nature, a transaction: it allows parties to close a dispute without obtaining a detailed verdict on all claims. Anthropic does not receive, in exchange for its payment, general legal validation of its past practices. Nor do the authors obtain a ruling establishing, point by point, the company’s liability for all works included.

The distinction is fundamental when interpreting the precedent. The case creates a financial and strategic precedent, but it does not by itself create exhaustive case law on training large language models. US courts will still have to answer several distinct questions: the status of lawfully acquired copies; the role of the transformation brought about by training; possible outputs reproducing protected passages; the effect on the market for original works; and liability related to the creation of datasets.

In Judge Alsup’s earlier decision, the distinction between training on lawfully acquired books and using pirated libraries was already explicit. The settlement does not erase that distinction. On the contrary, it underscores its practical importance. For a lab, the strongest line of defense is no longer simply to explain that the model transforms texts into statistical parameters. It also requires being able to demonstrate the origin of the data, the rights held, the contracts concluded, opt-out mechanisms and dataset traceability.

This traceability requirement runs up against the historical reality of generative AI. For a long time, competitive advantage was associated with the size of datasets and the speed of collection. Data available on the web, digitized archives and vast libraries of texts were seen as abundant raw material. Yet the Anthropic settlement is a reminder that raw material that is technically accessible is not necessarily legally usable without risk.

Fair use remains at the center of the debate, without being settled by the agreement

Fair use is often presented as a simple answer to the question of AI training, when it is in fact a case-by-case legal analysis. US law traditionally examines several factors, including the purpose and character of the use, the nature of the protected work, the amount used and the effect on the work’s potential market. The rise of generative AI places unprecedented pressure on these criteria, because systems often require the absorption of enormous collections of works, including entire creative works.

Technology companies emphasize the transformative nature of training. Their argument is that a model does not retain a searchable library in the same way as a download service or a search engine. It learns statistical relationships in the data in order to generate new responses. Under this interpretation, the purpose would differ from that of a reader seeking access to the text of a novel or an essay.

Authors and rights holders put forward different reasoning. For them, the scale of collection, the commercial exploitation of models and the existence of a potential licensing market must be taken into account. They also stress that generative systems can compete with certain forms of creation, writing, illustration, translation or summarization, while having been developed using works that were neither authorized nor compensated.

The Anthropic case highlighted these tensions, but its settlement does not resolve them for the entire sector. The earlier decision favorable to Anthropic on lawfully acquired books is important because it suggests that a court may regard training as transformative. However, it is not a general permission to use every available work. On the one hand, it is tied to the facts of the case. On the other, other US jurisdictions may assess the circumstances differently, particularly depending on the type of content, the conditions of access to the data, the model’s capabilities or its effect on markets.

Several stages that are sometimes conflated in public debate must also be distinguished. The first is acquisition: did a company obtain a copy lawfully? The second is technical reproduction, because training often involves copying, cleaning, converting and storing content. The third is training itself. The fourth is the system’s output: can it reproduce protected passages or offer an alternative to the original content? Finally, there is the commercialization of the product, its subscriptions, its APIs and its enterprise uses.

A settlement can resolve certain risks associated with a given stage without clarifying all the others. This is precisely what makes the Anthropic case so important: it shows that a player can have serious arguments regarding the transformative nature of training while still being exposed to a considerable bill over the provenance of the copies used.

This situation also differs from debates concerning content directly posted online by publishers, media outlets or creators. In such cases, public access to a web page does not necessarily mean that mass extraction of its content to train a model is authorized. Likewise, the existence of terms of use, opt-out files or contractual restrictions may have a significance distinct from copyright itself. The settlement approved in the Anthropic case does not decide these issues either.

For labs, the legal message is therefore twofold. The fair use argument retains a place in their defense strategy, particularly for lawfully obtained data. But it can no longer be viewed as a universal solution allowing them to ignore the history of the data supply chain. The greater a model’s economic value, the more closely the conditions under which its corpus was assembled are likely to be examined.

A financial signal for Anthropic and its competitors

The settlement represents a very high cost, but its significance is also measured by what it tells other AI players. Cutting-edge models require considerable investments in computing power, data centers, energy, research and recruitment. This economic equation now more visibly includes the potential cost of data: licenses, provenance checks, litigation, insurance, provisions and negotiations with rights holders.

For Anthropic, the case comes as the company has become one of the leading private generative AI groups. Amazon and Google have invested in the company, while its Claude models are used through products aimed at the general public and businesses. In such a context, a $1.5 billion settlement is both an exceptional charge and a demonstration of the risk’s maturity: content-related litigation is no longer peripheral to technological competition.

Competitors are necessarily watching this decision. OpenAI faces several lawsuits related to training data, notably from authors and media organizations. Meta has also been targeted by complaints over the use of books and other content. Google, Microsoft, Stability AI, Midjourney and other companies in the sector have faced similar challenges, under different legal configurations. Each case depends on its own facts, the works involved and the court hearing it; it would therefore be unwise to treat the Anthropic settlement as an automatic scale applicable to all proceedings.

But the amount acts as a benchmark. When a lab evaluates the cost of a license, it can no longer merely compare that expense with a hypothetical saving from collection. It must also consider the risk of a class action, uncertainty over damages, executives’ lost time, restrictions potentially imposed by courts and the impact on relationships with business customers.

The issue is particularly sensitive for companies selling AI tools to large organizations. Banks, insurers, healthcare groups, public administrations, law firms and publishers want to know where the data used by the systems they deploy comes from. Their own regulatory or reputational exposure may depend on the guarantees offered by their suppliers. A customer company does not necessarily control the training corpus of a general-purpose model, but it can request contractual commitments, indemnification mechanisms or offerings using licensed data.

The settlement could therefore accelerate market segmentation. On one side, some general-purpose models will seek to defend the lawfulness of very large corpora, relying on fair use and favorable decisions obtained in certain cases. On the other, offerings aimed at the most sensitive enterprises may place greater emphasis on controlled data, partnerships with publishers, proprietary corpora or mechanisms for training on customers’ internal data.

This development does not mean that licensing will become the only possible path for AI. Ongoing litigation in the United States will continue to determine the scope of fair use. However, the Anthropic agreement makes it harder to maintain the idea that access to an immense pool of content, even if collected without authorization, can be treated as a technical detail that can be fixed afterward.

The precedent also concerns investors. Until now, AI lab valuations have often been based on model performance, access to computing, commercial growth and their ability to attract partners. The implicit legal debt of datasets may now appear more clearly in due diligence. A company that cannot explain the origin of certain corpora precisely may represent a risk whose scale extends far beyond ordinary legal costs.

Direct implications for Europe and the French-speaking market

The Anthropic settlement is American and rests on US law. It therefore cannot be mechanically transposed to France or the European Union. The notion of fair use does not exist in French law in the same form. The European framework is based on more clearly defined copyright exceptions, including text and data mining exceptions, often referred to by the English expression text and data mining, or TDM.

The European directive on copyright in the digital single market introduced two important mechanisms. One benefits research organizations and cultural heritage institutions, under certain conditions. The other allows text and data mining for other players, including commercial ones, provided that the content is lawfully accessible and rights holders have not expressly reserved their rights.

This European architecture makes data provenance particularly important. Lawful access to content is a central condition. Rights reservations expressed by rights holders may also limit use under the commercial exception. In France, publishers, news agencies, media outlets and collective management organizations are therefore closely monitoring the technical and contractual mechanisms through which AI developers collect online content.

The US case does not answer all European questions, but it reinforces an idea already present in debates in Brussels and Paris: generative AI will have to contend with transparency obligations and the economic interests of creators. The European Artificial Intelligence Act, the AI Act, notably provides for specific obligations for providers of general-purpose AI models, including the development of a policy to comply with EU copyright law and the publication of a sufficiently detailed summary of the content used for training.

These obligations do not automatically create a compulsory license for all content. Nor do they say that a provider must publicly disclose every file included in its corpus. But they increase pressure in favor of documented data governance. In light of the Anthropic settlement, the logic becomes more tangible: transparency is not merely an ethical debate. It can determine a company’s ability to demonstrate that it has respected applicable rights.

For French-speaking players, the consequences are multiple. Book publishers, press groups, audiovisual producers, cultural platforms and independent creators can view this precedent as an additional argument in favor of licensing negotiations. French and European startups, meanwhile, must incorporate the issue of the rights chain very early into their data choices, including when they build specialized models rather than very large general-purpose models.

This constraint may appear heavier for local players with fewer resources than US giants. Yet it can also become a differentiating factor. A model specialized in French, legal, medical, scientific or industrial corpora, obtained under a clear contractual framework, can meet the demand of organizations that prioritize compliance and data sovereignty over model size alone.

France also has a dense publishing and cultural ecosystem, in which copyright holds significant economic and political importance. Negotiations between AI developers and catalog holders could therefore become an industrial issue as much as a legal one. The Anthropic case does not set the price of a work, or the price of a training license. It does, however, remind us that ignoring this negotiation can cost far more than anticipated.

Toward a new corpus economy, between licenses, audits and case law

The next stage will play out less in the spectacular announcement of the settlement than in the operational changes it will encourage. Labs will have to better distinguish data from the public domain, content under open licenses, data accessible with rights holders’ permission, data covered by contracts, data subject to a rights reservation and collections whose provenance remains uncertain.

This mapping is not trivial. Training corpora have often been assembled over the years from multiple sources, with intermediate copies, deduplication operations, format conversions and third-party suppliers. For models already trained, tracing the entire chain can be particularly complex. Yet the Anthropic case suggests that a lack of documentation is no longer merely a governance shortcoming: it can become the weak point of a defense before courts or in collective negotiations.

Licensing agreements could take varied forms. Some may cover entire catalogs, others limited uses, time windows, content reserved for research or specific models. Publishers may seek fixed compensation mechanisms, payments indexed to use, transparency obligations or guarantees against the reproduction of substantial passages. Developers, for their part, will seek to preserve the ability to train high-performing systems without making their cost structure unsustainable.

The Anthropic settlement does not guarantee that this licensing market will develop uniformly. Rights holders do not all have the same bargaining power, works do not have the same value for models and uses are not alike. Individual authors, small publishers and regional media outlets could have greater difficulty reaching direct deals with major labs. This is one reason why class actions and rights management organizations will likely continue to play an important role.

US case law will also remain decisive. Other cases will have to clarify the relationship between training and fair use, particularly when data was not obtained from clearly pirate sources. Courts will also examine situations in which generative tools produce content too close to identifiable works, as well as the effects of these systems on existing or emerging licensing markets.

For Anthropic, the $1.5 billion payment closes a particularly difficult chapter, but it does not end the broader debate over Claude, large language models or the responsibilities of the companies developing them. For other labs, the lesson is broader: the law governing training data is becoming a lasting component of technological competition.

In the long term, the question will probably not be one of choosing between AI trained on the entire web and AI trained only on licensed content. Rather, the market is likely to be structured around a mix of public data, authorized data, content under contract and legal defenses that vary according to use. Approval of the Anthropic settlement does not provide a definitive answer on fair use, but it makes one thing much clearer: in the generative AI economy, the origin of data is set to matter as much as its volume.

Back to all news

Comments· No comments yet

Be the first to react.

Leave a comment