business 5 min read

Anthropic Accuses China of Stealing Claude’s Brain

Anthropic says Chinese firms Alibaba and Moonshot AI ran massive distillation campaigns against Claude, pulling out its reasoning ability through hundreds of millions of disguised queries. The revelation forces a reckoning: if closed models can be hollowed out, open-weight models may be even more vulnerable.

  • Alibaba
  • Anthropic
  • Open Weight Models
  • AI & Security
  • Moonshot AI
  • US-China Tech
  • AI Distillation

The Theft Wasn’t in the Code. It Was in the Thinking.

Anthropic’s latest threat intelligence report does not read like a routine security disclosure. It reads like a breach of something far more valuable than proprietary weights: the distilled reasoning patterns that make Claude distinct. The company confirmed five large-scale distillation campaigns linked to Chinese organizations, involving roughly 200 million Claude conversations. The two names on the list—Alibaba and Moonshot AI—signal that this was not the work of hobbyist researchers. This was industrial-scale knowledge extraction.

Distillation attacks work by feeding a powerful model carefully crafted prompts designed to coax it into revealing its internal reasoning. Anthropic already limits what users see: instead of raw chain-of-thought tokens, Claude surfaces only a summarized thought process. The attackers bypassed that guardrail by disguising their requests. One confirmed prompt asked Claude to translate prior work memory into natural, accurate Japanese katakana. The goal was not translation. The goal was to get Claude to output the very reasoning traces Anthropic had deliberately compressed.

Alibaba’s Campaign Was Remarkably Systematic

The largest operation, tied to Alibaba, generated an estimated 151 million conversations between May and July. That is roughly 3 million interactions per day across 3,500 accounts. The uniformity of the fixed prompts used across those accounts led Anthropic to conclude this was a single coordinated campaign aimed at harvesting training data for Alibaba’s Qwen model family. Qwen is already one of the most capable open-weight model series in the world. If Anthropic’s assessment holds, Qwen’s recent performance gains may rest partly on distilled Claude reasoning—extracted without license, without payment, and without attribution.

That framing matters enormously. The AI industry has spent years debating whether open-weight models accelerate innovation or erode the economic incentives for building them. This incident shifts the question from philosophical to legal. If the most advanced closed models can be hollowed out at scale, the economic case for keeping weights locked down grows stronger—and the case for open-weight models becomes harder to justify to anyone footing the compute bill.

The Military Angle Is Disturbing

The Moonshot AI campaign, behind the Kimi chatbot, carried a different signature. Anthropic found requests asking Claude to analyze CCTV footage for abnormal human behavior. The phrasing and intent pointed to Chinese military origins. Over ten days, 5,000 accounts funneled roughly 300,000 requests at Claude’s most capable Opus model. The target selection is telling: Opus is the version with the longest context window and the strongest reasoning. You do not attack Opus unless you want the highest-quality distillation possible.

This is the first time a distillation campaign has been explicitly linked to military intelligence activity against a US AI model. The precedent it sets could reshape export controls and AI-related sanctions within months.

Why This Changes the Open-Weights Debate

Open-weights advocates have long argued that sharing model checkpoints accelerates research, improves safety through scrutiny, and prevents a handful of well-funded labs from controlling the frontier. None of those arguments assume the weights themselves will be stolen through prompt engineering. But the boundary between “open weights” and “extracted capability” is thinner than the community likes to admit. If a company can distill Claude’s reasoning patterns without ever seeing Claude’s weights, then the value of open weights diminishes while the incentive to leak or exfiltrate them increases.

The practical implication is stark: open-weight models may become less defensible as a policy position, not because they are unsafe in the traditional sense, but because they offer less protective value in an environment where capability can be siphoned through interaction alone. A company that open-weights its model faces two vectors of exposure: direct weight theft and indirect capability extraction. Anthropic’s report makes clear that the second vector is now operational at nation-state scale.

What Comes Next

Anthropic has not released forensic proof of ownership for the accounts behind these campaigns. The linkage to Alibaba and Moonshot AI is based on behavioral patterns, shared infrastructure signals, and prompt uniformity. Both companies have not publicly responded. The allegations will likely face legal and evidentiary challenges before any regulatory body.

But the policy trajectory is already visible. The US Treasury and Commerce Department have been moving toward tighter controls on advanced AI model access by foreign entities, particularly Chinese ones. This report provides concrete documentation to support those efforts. Expect proposed rules that require US cloud providers and API operators to detect and block distillation-style abuse patterns, and to flag accounts that exhibit the kind of automated, high-volume reasoning extraction seen here.

The broader industry should also take note. OpenAI reported in February that DeepSeek had conducted similar distillation activity against GPT. Anthropic disclosed a Chinese research institute attack earlier this year. The pattern is consistent: Chinese labs and companies are converging on distillation as a shortcut to frontier capability. The US response will not be purely defensive. It will reshape how AI models are distributed, monitored, and regulated globally.

For developers building on top of frontier models, the message is straightforward: your reasoning is now an asset under siege. The contracts, the terms of service, and the technical safeguards around API access will harden. For the open-weights community, the challenge is sharper: prove that openness survives in an environment where capability can be extracted without ever opening the box.