Anthropic accuses three major Chinese artificial intelligence players — Alibaba, Moonshot AI and DeepSeek — of conducting repeated distillation campaigns against its models. The information, reported by TechCrunch in an article titled “Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek”, goes beyond a mere dispute between companies. It raises a question that has become central to the economics of large language models: at what point does using the responses produced by a competing AI cease to be a form of learning or technical comparison and become abusive appropriation of know-how?
According to Anthropic, the attempts attributed to Alibaba, Moonshot AI and DeepSeek have intensified in recent months. The US company describes these activities as distillation campaigns, meaning the systematic exploitation of a cutting-edge model’s outputs to train, improve or evaluate another model. Anthropic is therefore not merely describing occasional use of its Claude assistant by developers or end users. The complaint concerns a structured, large-scale approach likely to turn the observable capabilities of a proprietary system into training data for a competitor.
The companies cited by Anthropic occupy different but important positions in China’s generative AI ecosystem. Alibaba develops the Qwen family of models, distributed in open and commercial variants. Moonshot AI became known in particular through the Kimi assistant. DeepSeek became one of the industry’s most closely watched names after releasing reasoning models and open models that generated strong international interest. None of these three companies is presented here as guilty by a court: Anthropic’s report is a corporate accusation, relayed by TechCrunch, rather than a judicial or administrative decision.
The stakes are nevertheless considerable. The best language models do not rely solely on parameters, chips and massive corpora. They also embody years of data selection, training, alignment, post-training tuning, safety assessments and interface optimization. When a model provides a detailed, structured response tailored to a complex instruction, it can reveal part of this invisible work. Distillation aims precisely to turn these observable behaviors into useful material for another system.
Distillation: a legitimate technique that has become a battleground
Distillation is not, in itself, a fraudulent practice. It has long been part of the machine learning toolkit. In its conventional sense, distillation involves using a teacher model, generally larger or more expensive, to train a more compact student model. The student model seeks to reproduce certain behaviors of the teacher model: its predictions, the way it classifies examples or, in the case of large language models, its responses to instructions.
This approach has historically been associated with a simple idea: a powerful model can transfer some of its capabilities to a lighter model. In a controlled setting, the same company can, for example, train a very large system, then use it to produce synthetic data intended for a faster, cheaper version or one adapted to a specific task. This method can reduce inference costs, enable operation on more modest infrastructure and facilitate the deployment of features in consumer products.
The issue raised by Anthropic lies elsewhere: the presumed teacher model would not be owned or controlled by the organization building the student model. It would be a proprietary system accessible through an interface, an API or a conversational product, with terms of use that normally govern automation, large-scale extraction and the use of outputs to train competing models.
In large language models, the technical boundary is particularly difficult to draw. A user may legitimately ask an assistant to summarize a text, write code, correct a translation or solve a logic problem. A laboratory may also compare several models on test sets to measure their accuracy, cost or resistance to certain instructions. Such uses fall under evaluation, research, software development or normal use of a service.
The situation changes when requests are designed, multiplied and archived in order to build a corpus of responses intended to reproduce the behaviors of a third-party model. This may involve very large volumes of instructions, varied automatically generated prompts, structured collection of outputs or attempts targeting specific capabilities: mathematical reasoning, code generation, multilingual writing, following complex instructions or refusing sensitive requests.
The distinction is important. A model does not merely reproduce a raw knowledge base. Its responses are the result of an architecture, prior training, human-preference tuning, safety mechanisms and product choices. Massively collecting these responses can therefore provide access to a form of teaching signal that would be costly to obtain otherwise. The value of that signal rises with the quality of the model being queried and the diversity of tasks to which it is subjected.
In this context, external distillation is sometimes described as a form of industrial shortcut. That expression must nevertheless be used with caution: it is not enough to establish wrongdoing, and the exact methods matter. A competing model does not automatically become a copy because it answers similar questions well. Large models are trained on immense data sets and can converge on comparable behaviors for standard tasks. But the organized exploitation of outputs from a closed system raises different issues, particularly contractual, technical and potentially commercial ones.
The issue is also linked to the rise of synthetic data. As high-quality public data become harder to collect, more protected by law or already widely exploited, the responses produced by good models themselves become a resource. They can be used to create reasoning examples, generate dialogues, produce programming exercises or rephrase content in many languages. The leading model then becomes not only a product, but also a highly sought-after data source.
What Anthropic claims about Alibaba, Moonshot AI and DeepSeek
In the report mentioned by TechCrunch, Anthropic details what it describes as distillation campaigns involving Alibaba, Moonshot AI and DeepSeek. The central point of the accusation is repetition: Anthropic says it identified attempts that did not amount to a few isolated interactions with Claude, but to operations intended to leverage its outputs for model development purposes.
The publication of such a report is itself significant. Companies operating closed models rarely communicate in detail about abusive behavior detected in their systems. They may fear revealing their monitoring methods, providing ways to evade them or triggering a diplomatic and commercial conflict. By making this accusation public, Anthropic instead chooses to make distillation a subject of open debate within the industry.
TechCrunch presents the case in a context of heightened technological competition between the United States and China. Anthropic links the activities it observed to three organizations developing their own models. This distinguishes the case from conventional abuses committed by spammers, fraud operators or users seeking only to circumvent access restrictions. The accusation targets industrial competitors that themselves possess significant research and deployment capabilities.
Alibaba is one of China’s largest technology groups and a major cloud player. Its Qwen models have helped establish the company among the platforms that matter in the distribution of language and multimodal models. The group has computing capabilities, a presence in cloud services and commercial outlets. Any accusation concerning it therefore takes on a dimension beyond the research laboratory: it concerns one of Asia’s major technology groups engaged in the race for AI infrastructure and applications.
Moonshot AI belongs to a newer generation of Chinese startups specializing in conversational assistants and language models. Kimi, its flagship product, has stood out in a highly competitive Chinese market, where companies seek to offer tools capable of processing long documents, assisting users in information retrieval and responding to complex queries. Pressure to improve the perceived quality of models quickly is particularly strong in this segment.
DeepSeek, lastly, occupies a distinctive place in the global debate on models. The company has drawn considerable attention through its publications and the release of open models or models available under licenses allowing their examination and deployment. Its work has fueled discussions on training costs, reasoning, model efficiency and the gap between closed US platforms and open Chinese alternatives. Anthropic’s accusation thus adds a sensitive dimension to a player already at the center of debates on the international diffusion of AI capabilities.
However, an essential distinction must be maintained between an allegation and its public demonstration. The brief reported by TechCrunch indicates that Anthropic accuses these companies of distillation campaigns; it does not turn that claim into a judicial finding. Without a court decision, full access to the technical evidence collected by Anthropic and adversarial publication of all relevant material, it would be improper to present these accusations as legally established facts.
This caution does not diminish the strategic significance of the case. Even when they do not lead to public proceedings, accusations of distillation can alter how AI providers manage access. They can lead to stricter controls on accounts, API keys, request volumes, usage patterns and attempts to create automated corpora. They can also fuel future litigation concerning terms of use, unfair competition, trade secrets or cross-border access to digital services.
Protecting a model without blocking research, evaluation and normal uses
The case illustrates a structural difficulty for model providers: an accessible system is necessarily observable to some extent. A company selling an API or a conversational assistant must enable its customers to submit requests and receive responses. Yet every response represents an example of the model’s behavior. The provider wants to make that behavior useful, but does not want it to be harvested at very large scale to train a competitor.
Protection therefore cannot rely solely on secrecy surrounding source code or parameters. In most cases, assistant users see neither the model weights nor the training infrastructure. But they can test its capabilities through thousands of questions. Protection then shifts to analyzing access behavior: request frequency and volume, repeated prompt structures, creation of multiple accounts, use of separate credentials, automation, traffic origin or correlations between seemingly separate activities.
This kind of detection is delicate. A player conducting robustness assessments, benchmarks or safety research may also send numerous unusual requests. Large corporate customers may legitimately test an API on thousands of documents. Research teams may compare several systems on sets of examinations, programming cases or multilingual scenarios. An overly aggressive defense mechanism could then penalize legitimate uses, reduce transparency or complicate access for independent researchers.
Providers nevertheless have several levers at their disposal. They can limit volumes, adjust quotas, require additional verification for certain access, identify automation patterns or strengthen clauses prohibiting the use of outputs to train competing models. These measures are more access-control tools than absolute protections. They do not eliminate the possibility of collecting responses, but they can make an operation costly, more visible and riskier.
This dynamic explains why distillation has become such a sensitive subject in the closed-model market. Companies invest considerable sums in computing infrastructure, recruiting researchers, acquiring data, annotation, evaluation and security. If competitors can recover a significant share of the value of that work by querying a final product, the economic incentive to fund the most expensive models may erode.
Conversely, it would be simplistic to equate all synthetic data generation or all functional imitation with illegitimate appropriation. The history of computing is made up of competing products that offer similar functions, follow common standards or seek to provide better performance. Language models are subject to the same reality: two systems trained on different data can learn to write code, summarize a contract or converse in French without it being possible to automatically infer a transfer of data between them.
The question therefore becomes evidentiary. What makes it possible to establish that a corpus was assembled from the responses of a specific model? How can competitive evaluation be distinguished from collection intended for training? Which technical indicators can be published without weakening the provider’s defenses? And what evidence would be admissible in countries whose rules on digital contracts, trade secrets and training data differ?
The answer will probably not be a single one. Companies will seek protection through their contracts and abuse detection. Courts, when called upon, will have to interpret often technical situations. Regulators may take an interest in effects on competition, system transparency and international data flows. Customers, meanwhile, will seek a different guarantee: that restrictions against distillation do not make APIs too closed, too closely monitored or too difficult to integrate into real products.
A new front in the technological rivalry between Washington and Beijing
The choice of the three targeted companies immediately gives the case a geopolitical dimension. Competition in AI is no longer only about chatbots and conversational interfaces. It concerns semiconductors, data centers, development tools, open models, export rules, investment and access to foreign markets. Distillation accusations add another angle: the indirect circulation of capabilities through the outputs of models accessible online.
For several years, the United States has strengthened its controls on exports of advanced computing technologies to China, particularly in the field of chips used for artificial intelligence. The stated aim of these measures is to limit access to certain cutting-edge computing capabilities. But the availability of a US model through an API creates a different channel for the potential transfer of capabilities: not the export of a chip or the delivery of model weights, but access to the results produced by that model.
This difference is essential. A textual response is not equivalent to delivery of the model itself. It does not, by itself, make it possible to reconstruct all the parameters, training data or architecture. However, repeated across a very large number of relevant examples, it can constitute useful supervision material. This is precisely what makes distillation difficult to fit into the conventional categories of export controls.
A sanctions or trade-restrictions regime is generally designed for identifiable goods, software, equipment or services. Remotely accessible models blur this separation. A provider can host its system in the United States or another country, charge for software access and serve users spread around the world. Determining who is actually using a service, on behalf of which organization and for what purpose is more complex than identifying the purchaser of physical hardware.
Anthropic’s accusations may therefore fuel calls for stronger control over access to the most advanced models. But that response carries risks. Broader closure of US services could accelerate demand for self-hosted models, open systems or providers established outside the United States. It could also further fragment the global market, with technical and regulatory ecosystems that are less interoperable.
The DeepSeek case makes this tension particularly visible. The rise of open or downloadable models has changed the balance of the sector. For developers, businesses and public administrations, access to a model that can be deployed locally offers greater control over data, costs and customization. For closed-model providers, this openness can reduce their ability to impose terms of use after distribution. Competition therefore does not merely pit the United States against China; it also pits economic models based on the controlled API against more open distribution strategies.
The rivalry may also have a paradoxical effect. The more leading providers lock down access to their models, the more they confirm that their models’ outputs represent a strategic resource. The more they claim those outputs can be used to improve competitors, the more they reinforce the idea that the observable behavior of an AI has independent economic value. That value then becomes an object of protection, just like data, algorithms or infrastructure.
For Anthropic, the sequence fits into a corporate identity centered on the safety of AI systems. Founded by former OpenAI members, the company has built its reputation around Claude and work on alignment, including the approach known as constitutional AI. Its positions on AI risks and the need to govern advanced models give particular weight to its report. Distillation is not presented solely as a commercial loss: it can also be perceived as a way to recover model behaviors without gaining access to the same level of controls, assessments or safety guarantees.
Concrete consequences for Europe and the French-speaking market
In France and the rest of Europe, the case should not be read as a distant confrontation between US and Chinese companies. European organizations depend heavily on models developed outside the continent, which they consume through APIs, assistants integrated into software or cloud offerings. They may also use open models, adapt them to their own data and offer specialized versions. The distinction between evaluation, adaptation and distillation is therefore directly relevant to French companies.
A French software publisher comparing several models to select the best one for its customer service, translation or programming assistance is engaging in a normal market practice. A company training an internal model on data generated at scale by a competing API must, on the other hand, examine the provider’s contractual clauses very closely. The difficulty is not solely one of applicable law: it also concerns internal governance. Product, data and legal teams must know where the synthetic data used in their pipelines come from.
The European regulation on artificial intelligence, the AI Act, does not by itself settle the question of distillation between companies. Its primary ambition concerns safety, transparency and management of risks associated with AI systems. It also provides a framework for general-purpose AI models. But disputes over the use of API outputs may also fall under contract law, competition law, trade secrets, copyright depending on the circumstances, or national rules.
For European players, this overlap of frameworks is a practical issue. Companies must comply with their providers’ terms, protect personal data and confidential information sent to models, document their internal uses and anticipate the question of audits. If major laboratories strengthen their detection mechanisms against distillation, European customers could see an increase in verification requests, volume limits or restrictions on certain use cases.
The question of digital sovereignty also appears in the background. France and Europe are seeking to develop their own computing capabilities, foundation models and alternatives to US and Chinese platforms. Companies such as Mistral AI have helped make European ambitions in language models more visible. In this context, the debate over distillation has a dual significance: it concerns protecting European laboratories’ investments, but also their ability to legally and fairly access the best global tools to evaluate their own work.
A market that is too closed could penalize small companies and research laboratories that lack the resources of large groups. A market without credible protections could, conversely, make it more difficult to fund costly models and favor players capable of collecting data at scale or mobilizing the largest infrastructure. Europe will therefore have to strike a balance between scientific openness, competition, protection of know-how and strategic autonomy.
French adds a specific dimension. Leading models are often evaluated first in English, even though their multilingual performance is improving. High-quality responses in French, in fields such as law, public administration, health, finance or education, have particular value for local developers. Building French-language synthetic data sets is therefore a real industrial issue. But here again, the origin of the data, authorizations for use and contractual terms cannot be treated as technical details.
Toward an economy of model outputs and technical evidence
The episode revealed by TechCrunch probably signals a lasting development: AI-related conflicts will increasingly concern model outputs, not only the data that go into their training. Until now, much of the public debate has focused on corpora ingested by systems: content published online, protected works, personal data, press archives, source code or databases. Distillation shifts the focus toward content generated after training.
This development is logical. The outputs of the best models condense capabilities that are difficult to reproduce. They can incorporate apparent reasoning, educational presentation, a dialogue style, programming strategies or calibrated refusals. Even if they do not provide access to the entire system, they can offer leverage for training smaller or more specialized models. Value then lies in the behavior delivered to the user, not only in the invisible assets hosted in a data center.
Providers will likely seek to improve the traceability of these behaviors. Research has long explored various forms of watermarking or statistical signatures for AI-generated content. In text, these solutions are difficult to make robust: a response can be rephrased, translated, summarized or mixed with other data. Their use for evidentiary purposes therefore remains complex. Convincing evidence will likely need to combine several elements: access logs, request structures, volumes, links between accounts, technical analyses and possibly internal material obtained as part of proceedings.
The battle will not be limited to technical protections. Contracts will become more explicit about reusing outputs, creating synthetic data and training models. Major corporate customers, for their part, will demand clear rules: can they use the results of an assistant to improve an internal search engine? To produce customer-support examples? To train a narrowly specialized classifier? The answers will depend on providers, offerings and licenses, which risks adding complexity to an already fragmented landscape.
In the long term, the case pitting Anthropic against the Chinese players it cites could accelerate the separation between several categories of services. On one side, highly capable closed models, offered under controlled access and accompanied by strong restrictions on the reuse of their outputs. On the other, open or locally deployable models, for which users assume more technical and legal responsibility. Between them, models under intermediate licenses, allowing certain commercial uses but imposing conditions on redistribution, training or volumes of use.
This segmentation will not entirely solve the problem. Companies will continue testing each other’s models, users will continue comparing their responses and the boundaries between evaluation data, synthetic data and training data will remain porous. But Anthropic’s publication marks a milestone: it turns a practice often discussed in technical circles into an explicit subject of industrial governance and international rivalry.
The future of competition between models will therefore depend less on an abstract ban on distillation than on the ability of stakeholders to define verifiable rules. If protections become too opaque, they could limit research and entrench dominant platforms. If they remain too weak, investments in closed models could become harder to defend. Between these two extremes, the challenge for US, Chinese and European laboratories will be to prove they can protect their capabilities without turning access to artificial intelligence into an entirely segmented market.
Comments· 1 comment
Thanks for the clear overview—this is a fascinating and important story to follow, especially as competition in AI becomes increasingly global.