Model distillation began as a clever way to make artificial intelligence smaller. It is rapidly becoming one of the industry’s biggest arguments over who owns intelligence once a machine can reproduce it.
The basic technique is legitimate and widely used. A powerful “teacher” model generates answers that train a smaller “student,” transferring some of its capabilities at far lower cost than starting again from raw data. AI labs use it to create faster, cheaper versions of their own systems.
Anthropic says the same method is now being used against frontier labs at industrial scale.
In February, the company accused DeepSeek, Moonshot AI and MiniMax of generating more than 16 million exchanges with Claude through roughly 24,000 fraudulent accounts. Anthropic said the campaigns focused on reasoning, coding, computer use and tool orchestration, and relied on proxy services that distributed requests across large networks of accounts. (Anthropic)
The allegations are Anthropic’s, not independently adjudicated findings. Chinese officials have rejected broader accusations of illicit distillation as groundless. The dispute also carries an obvious commercial and geopolitical dimension: US laboratories want to protect models that cost billions to build, while Chinese developers are competing under restrictions on advanced chips and some model access. (Associated Press)
Still, the security problem is bigger than a quarrel over terms of service.
“A distillation campaign looks less like someone copying a file and more like an account-fraud and data-exfiltration operation conducted through valid product interfaces,” said Iliya Fayans, a cybersecurity expert and former major in Israel’s elite Unit 8200. “The individual prompts can look harmless. Detection depends on connecting identity, payment, infrastructure and behavioural signals across thousands of accounts.”
Anthropic says one proxy network controlled more than 20,000 accounts simultaneously, mixing suspected extraction traffic with ordinary customer requests. It has responded with behavioural classifiers, stronger verification, detection for coordinated accounts and systems intended to identify attempts to elicit reasoning traces.
The deeper question is whether a model’s behaviour can be protected in the same way as its weights or source code. Outputs are the product being sold. Customers must be allowed to query a system and often use its responses to improve their own software. A rule that blocks every repetitive or technical workload would punish exactly the developers an API is designed to attract.
That makes distillation defense an unusually delicate security discipline: the provider must identify intent from patterns without making the product unusable.
It also ensures the fight will not end at the API gateway. As open models improve, the value of high-quality reasoning data rises. As frontier models become better teachers, competitors need fewer original breakthroughs to build capable students.
The labs may keep their weights closed. Their knowledge is already leaving one answer at a time.





