technology 7 min read

Anthropic Fires the First Shot in the AI Distillation War

Anthropic's September 2026 misuse report names Chinese labs in an industrial-scale campaign to steal Claude's capabilities through distillation — but the real story is what the data reveals about the asymmetry between Western AI developers and state-linked Chinese competitors who treat model weights as open-source.

  • Anthropic
  • AI Safety
  • Chinese AI
  • DeepSeek
  • AI Governance

The report that came a day too late

Anthropic published its September 2026 threat intelligence report on September 10. The day before, a former Anthropic employee named Jacob Coxson posted on X that both OpenAI and Anthropic had failed to act responsibly. The post cited Evan Hubinger and Samuel Marks — two senior Anthropic figures — for statements that AI could cause human extinction. The timing was not accidental. An IPO is expected this year. A report that catalogues every way your product can be weaponized, while your own staff publicly questions whether you are doing enough, is a hard look in the mirror.

The report itself, titled “Detecting and countering misuse of AI: September 2026,” covers cases from December 2025 through August 2026. It organizes abuse into seven categories: cyberattacks, influence operations, surveillance, fraud, biological misuse, conventional weapons development, and — most newsworthy — distillation. That last category is where the report breaks most clearly from prior AI safety literature, and where the geopolitical stakes sharpen into something you can point at on a map.

Distillation as industrial espionage

The report names five Chinese AI companies by name: Alibaba, DeepSeek, Moonshot AI, MiniMax, and Z.ai. It also names Xiaomi. These firms, according to Anthropic’s analysis, ran an “industrial-scale covert campaign” to extract Claude’s capabilities using fake accounts, stolen credit cards, and compromised authentication credentials. The technique is called distillation: you feed another model’s outputs into your own training pipeline until it approximates the target’s behavior, without ever touching the original weights.

The methodology is not secret. It has been discussed in academic circles for years. What made this campaign notable was its scale and its targeting. Anthropic says it intercepted more than 16 million requests — many routed through third-party model-routing services that Western developers use daily. The requests came from over 1,500 proxy accounts. They carried conversations, code snippets, and developer environment reconstructions designed to teach a replacement model how Claude thinks.

Xiaomi’s case is the most detailed. According to Anthropic, the company’s internal model MiMo did not handle the traffic itself. Instead, it silently forwarded customer requests to Claude, captured the responses, and used them to generate supervised fine-tuning and reinforcement-learning data. The captured sessions contained the names and contact details of hundreds of Xiaomi users, corporate data in more than a dozen languages, and development environments reconstructed turn by turn. Claude’s responses were not relayed back to the original requesters. They were archived, processed, and fed into Xiaomi’s training loop.

Moonshot AI received similar treatment. The report alleges that instead of processing requests through its own Kimi model, Moonshot silently forwarded them to Claude. This is a classic distillation pattern: let the target model do the work, learn from its answers, never reveal that you are learning.

Why distillation matters more than theft

There is a difference between stealing model weights and stealing model behavior. Weights are static. They can be patched, rotated, replaced. Behavior is harder to defend. Once a distillation campaign succeeds, the resulting model can operate inside the same legal and jurisdictional boundaries as any domestically-trained alternative. It can be deployed in enterprise products, integrated into government systems, sold to third parties — all without a single byte of stolen intellectual property crossing a border.

This is why the Chinese lab campaign is structurally different from earlier AI safety concerns. The old worry was that a rogue actor would breach a model provider’s infrastructure and walk out with weights. The new worry is that legitimate users, acting as unwitting data sources, will train a rival model through billions of carefully constructed interactions. The breach is not a single event. It is a slow leak across thousands of accounts, each one looking like normal usage.

Anthropic’s report notes that Claude Haiku, Sonnet, and Opus were all exploited. The next-generation Fable and Mythos models saw only one confirmed distillation attempt — suggesting that either the newer architectures are harder to distill, or that the campaign has not yet scaled to them. That distinction matters for investors watching Anthropic’s IPO timeline.

Biological misuse: the ambiguity problem

If distillation is the report’s geopolitical centerpiece, biological misuse is its moral one. Anthropic documented five cases involving high-consequence biological research: gain-of-function studies on avian influenza, peptide database construction paired with toxicity-optimization pipelines, and other work that sits in the gray zone between legitimate vaccine research and pathogen design.

The difficulty is methodological. The techniques used to design a safer vaccine and the techniques used to design a deadlier pathogen overlap significantly. Anthropic says it could not determine intent in several cases. It chose to block on the basis of capability risk rather than proven恶意 — a decision that raises uncomfortable questions about who gets to decide when scientific research crosses a line.

The report notes that older Claude models like Opus 4 and Sonnet 4.5 were not capable enough to meaningfully assist with advanced biological research, and that safeguards on those models were limited. Newer generations have tighter restrictions. But the asymmetry is stark: the same capability threshold that makes a model dangerous for biological misuse also makes it valuable for distillation. The firms running the distillation campaigns are not trying to design pathogens. They are trying to build models that can — and eventually will.

Surveillance and influence: the state-linked pattern

The report’s surveillance section describes accounts linked to Chinese state security organs using Claude to monitor and suppress dissent. Targets included Hong Kong pro-democracy activists, Tiananmen commemoration participants, Uyghur support networks, and Western human rights organizations. One tracked account, @whyyoutouzhele, had approximately 2.1 million followers on X and posted content that included videos of domestic protests and censored information relayed from within China.

Influence operations covered nine cases. Russian state media outlets were found using Claude in their article and broadcast script production pipelines. Three Iranian state-linked propaganda organizations appeared in the data. A Malaysian election was targeted by a Turkish firm selling an opinion-manipulation platform. A French advertising company operated approximately 70 fake news sites publishing roughly 8,900 articles across them.

The pattern is consistent: state and state-adjacent actors are integrating Claude into existing operational workflows the same way they integrate other commercial software. The difference is that Claude’s capabilities are harder to audit than traditional tools. When a government uses a word processor to draft a propaganda piece, you can trace the document. When it uses a language model to generate content at scale, the provenance is fuzzy.

What happens next

The report arrives at a moment when the AI industry is still defining its relationship with government. Anthropic’s IPO, expected this year, will force the company to disclose risk factors that include its own product being used against the interests of the countries where it operates. The distillation campaign, if Anthropic’s characterization is correct, suggests that Chinese AI firms are treating Claude not as a product to compete with but as a dataset to consume.

The response options are limited. Distillation-resistant model design is an active area of research, but no solution is proven at scale. Usage monitoring can catch anomalies, but the proxy-and-rotating-account technique described in the report is designed to look like normal traffic. Legal recourse across jurisdictions is slow and uncertain. The most immediate effect may be reputational: any organization caught running a distillation campaign faces the same kind of public naming that Anthropic has now applied to Alibaba, DeepSeek, Xiaomi, and the others.

The Coxson post from the day before the report was published is worth reading alongside the document itself. It raises a question that the report does not fully answer: if Anthropic and OpenAI are both failing to act responsibly, what standard are they being measured against? The distillation campaign is one failure mode. The biological misuse ambiguity is another. The surveillance integration is a third. Each one points to a different kind of responsibility — technical, ethical, and geopolitical — and none of them has a clean solution.

What is clear is that the AI safety conversation has moved past theoretical risk assessment into documented, attributed, named-and-shamed reality. The September 2026 report is not a warning about what might happen. It is a record of what is happening. The distillation war is already underway, and Anthropic has just published the first public casualty list.