GREG BAKER

Anthropic's latest allegation over model distillation, submitted to the US Senate in June, named Alibaba's Qwen team in what the company called the largest campaign of its kind in its history. It said 28.8 million exchanges had been extracted through roughly 25,000 fake accounts over six weeks. A February report had accused DeepSeek, Moonshot AI and MiniMax of similar but smaller-scale activity. Alibaba's Hong Kong-listed shares fell as much as 5% during the day to a four-month low of HK$94.55. None of the four companies responded to requests for comment.

That is the immediate trigger. It is also close to beside the point.

The more important question is not whether Alibaba did what Anthropic alleges. That remains a legal dispute, with at least one Beijing intellectual property lawyer arguing that the case rests on nothing more than the scraping of publicly available API outputs. It will be resolved in Congress or in court, not in this column. The real question is why an entire tier of Chinese AI companies keeps turning to the same shortcut, openly enough that an industry research note estimated more than 60% of independent Chinese model teams were relying on distillation as of 2024. The figure is widely known within the industry, not a secret uncovered by Anthropic. So what does it actually reveal about the weaknesses of China's AI sector?

It is not a lack of originality. The simplistic narrative that China copies while the United States innovates does not stand up to the technical record. DeepSeek's multi-head latent attention mechanism and mixture-of-experts architecture were developed independently, reviewed in international journals and recognized as genuine contributions even by people skeptical of the company's broader claims. If the weakness were conceptual, distillation would not help: it is impossible to copy an idea without understanding it well enough to reproduce it. Teams that distill models still need to know what they are extracting and how to integrate it. The bottleneck lies elsewhere.

The constraint is frontier-scale computing, and it is structural rather than simply a matter of chip restrictions. Training a model large enough to produce genuinely new capabilities requires a volume of GPU hours that only a handful of Chinese companies can finance, and only a fraction of them can obtain hardware approaching Nvidia's top tier. China's exposure to Washington's chip controls—and its reliance on the deliberately downgraded H20 as a workaround—was tested earlier this year, when a brief supply disruption exposed how limited that substitute really was. But the deeper problem is not the chip ban itself. Pretraining a frontier model from scratch requires sustained, coordinated investment in computing on a scale that most companies, in any country, cannot easily justify given the risk of failure. Distillation is what a company turns to when it wants frontier-level output without taking on frontier-level risk. It is a rational response to a computing ceiling, not evidence of an idea deficit.

That is also why the industry is consolidating—and why consolidation is the real diagnostic. The same research note that put reliance on distillation above 60% also tracked the number of independent large-model companies in China falling from a peak of 237 to 112, with a further decline to below 50 projected. This is not a story about weak companies being caught stealing. It is a story about a two-tier structure: a small group of firms—DeepSeek, Alibaba and a few others—with the balance sheets and research depth to conduct genuine frontier pretraining, and a much larger group that never had a realistic path to that tier and used distillation as the only way to remain competitive on benchmarks in the meantime. Anthropic's crackdown, along with likely restrictions on export APIs, will not hit the first group particularly hard. It will remove the only lifeline available to the second.

What distillation does not transfer is the part that will matter most in the future. A distilled model inherits its teacher's surface behavior—fluent reasoning traces, coding styles and tool-use patterns—without inheriting the safety-alignment work, evaluation methods or understanding of failure modes that the teacher's own lab may have spent years developing. This is a real and underdiscussed asymmetry: benchmark scores can converge even as the underlying gap in research maturity remains wide, because benchmarks measure output quality rather than the depth of the process that enables a model to self-correct, refuse unsafe requests or generalize beyond its training distribution. If Chinese labs lose API-level access to frontier Western models—as both Anthropic's product restrictions and the proposed Hagerty-Kim sanctions appear to indicate—the companies that never built that underlying research capacity independently will be the first to plateau, not firms such as DeepSeek that already have it.

So the honest answer to the question of China's AI sector's real vulnerability is neither algorithms nor even chips in the narrow sense. It is that the industry's middle tier built its competitiveness on borrowed capabilities rather than proprietary research infrastructure, while every policy lever now being pulled in Washington—export controls, API access restrictions and, now, enforcement against distillation—is aimed precisely at that borrowed layer. The firms with genuine pretraini

Originally published on IBTimes Hong Kong