For years, large language models were trained on the open web. News articles, blogs, code repositories, academic papers, and creative writing were absorbed into vast training datasets. AI companies argued that this process was transformative, analytical, and comparable to how humans learn from reading.
Now, something unusual is happening.
Instead of training on human-created content, AI labs are increasingly accused of training on each other.
In February 2026,
Anthropic publicly stated that it had detected large-scale “distillation attacks” against its Claude model. According to the company, multiple AI laboratories used thousands of coordinated accounts to generate millions of interactions with Claude. The suspected goal was not casual use. It was a capability extraction systematically prompting the model in order to train rival systems on its outputs.
The method is known as model distillation. In technical terms, distillation allows a smaller or less advanced model to learn from the outputs of a stronger “teacher” model. The student does not need access to the teacher’s internal weights or original training data. It only needs enough high-quality responses to approximate the teacher’s behaviour.
Distillation itself is not new. It has long been used within companies to compress large models into smaller ones. What makes this situation different is the alleged scale and intent: external labs using a frontier model as a training oracle to accelerate their own development.
This raises an uncomfortable question for the AI industry:
Are large language models now receiving a version of the same treatment that helped create them?
From Web Scraping to Model Scraping
The early expansion of generative AI relied heavily on scraping publicly accessible internet content. Many creators, journalists, and artists later discovered that their work had been included in training datasets without explicit permission. Lawsuits followed. Debates about copyright, consent, and fair use intensified.
AI developers maintained that their models did not “copy” content in a conventional sense. They learned patterns. They extracted statistical relationships. They transformed data into mathematical weights.
But the distillation challenges that framing.
When one AI system repeatedly queries another to replicate its reasoning style, problem-solving structure, or coding ability, the distinction between learning and copying becomes less comfortable. If scraping the web was described as learning from publicly available information, is scraping a competitor’s model outputs any different?
Anthropic has characterised the alleged activity as a threat not only to commercial advantage but also to
AI security. The company argues that large-scale extraction campaigns undermine investments in computing, safety research, and engineering.
In other words, the “data moat” that once protected frontier labs may not be as secure as expected.
The New Competitive Battlefield
This shift signals a deeper transformation in AI competition.
The race is no longer only about who can collect more human data. It is about who can defend their model from becoming someone else’s training dataset.
Companies are responding in several ways:
- Detecting coordinated prompting patterns designed to extract reasoning traces
- Limiting automated access and suspicious account clusters
- Reducing the exposure of step-by-step reasoning in outputs
- Tightening verification pathways for high-volume usage
The frontier model is gradually becoming more guarded, less transparent, more monitored, and more defensive.
The irony is striking. An industry built on open-web ingestion is now building digital fences around its own outputs.
Where the Law Stands
This conflict is unfolding while regulators are still trying to define acceptable AI data practices.
In the European Union,
the Digital Single Market Directive allows text and data mining (TDM) of lawfully accessible works for commercial purposes unless rights holders explicitly opt out. This opt-out mechanism can be expressed through machine-readable signals.
However, many legal scholars argue that generative AI may go beyond traditional text and data mining. Unlike analytical extraction of facts, generative systems can produce expressive outputs that resemble original creative works. Whether this violates copyright law is still being tested in courts.
The newer
EU Artificial Intelligence Act adds another layer. Providers of general-purpose AI models must publish summaries of training data and implement policies that respect EU copyright law, including recognition of opt-out signals.
Notably, the law focuses on human rights holders, not AI models.
There is currently no legal category for “model-to-model extraction.” If one AI system uses outputs from another for training, the dispute falls into contract law, intellectual property law, or trade secret protection, depending on circumstances.
The regulatory system was designed to manage human content rights, not inter-model competition.
National Security and AI Sovereignty
The debate is not purely commercial.
Some U.S. policymakers argue that limiting access to advanced AI chips and frontier models is necessary to preserve technological advantage. If distillation allows competitors to close capability gaps without equivalent compute investment, export controls may be less effective than anticipated.
In this framing, model outputs become strategic assets.
But this also reveals a tension. If knowledge extraction is unacceptable when directed at frontier models, how should we view the original extraction of the web’s knowledge to build those models?
The industry’s narrative is evolving. What was once defended as transformative learning is now described, in some cases, as unauthorised capability harvesting.
Copy-Paste or Competitive Evolution?
It would be too simplistic to reduce this to “copy-paste.”
Distillation does not create an identical replica. It creates an approximation. The student model still requires engineering, training infrastructure, and refinement. It cannot simply duplicate weights.
At the same time, large-scale extraction campaigns blur ethical boundaries. If one company invests hundreds of millions in safety tuning and reasoning optimisation, systematic harvesting of those outputs may undermine incentives for innovation.
The deeper issue may not be copying. It may be symmetry.
When AI companies trained on human content, they framed it as technological progress. When AI companies train on AI outputs, the same logic suddenly appears more predatory.
The ecosystem is confronting its own mirror.
Toward a More Sustainable Model
If AI development continues as a cycle of extraction first from humans, then from other models, the web risks becoming closed and defensive. Publishers are restricting crawler access. Artists are deploying adversarial tools. AI companies are tightening APIs.
An extractive equilibrium is unstable.
A more sustainable path may involve clearer licensing systems, revenue-sharing models, and enforceable transparency requirements. The EU AI Act moves in that direction by mandating training data summaries and copyright policies. Whether it succeeds remains uncertain.
What is clear is this: the era of unchallenged data harvesting is ending.
When large language models begin scraping one another, the debate shifts from innovation versus regulation to something more fundamental: fairness in the digital knowledge economy.
The AI industry is no longer just negotiating with creators. It is negotiating with itself.
And in that negotiation, it may finally confront a difficult question:
Is this competitive evolution or simply the natural consequence of building intelligence on extraction?