America’s Open-Model Paradox
To Distill, or Not to Distill?

To Distill, or Not to Distill?
China increasingly supplies the models Western companies use to serve, train, and build AI.
Qwen’s share of new open-model fine-tunes and adaptations rose from 1% in January 2024 to 69% by February 2026 according to ATOM’s Report. The majority of American AI startups seem to be using Chinese open weights somewhere in their stack.
This dependence now extends upstream. Western application companies are building on Chinese Open Source models. Further, Western labs are using Chinese models as teachers and sources of synthetic training data in the torrid race to close the frontier gap.
For example, Thinking Machines pre-trained Inkling independently, but used synthetic data generated by Moonshot’s Kimi K2.5 to bootstrap its supervised fine-tuning.
The relevant point is not how much of Inkling came from Kimi. The point is that a Western lab had a legal path to learn from a Chinese open model, while equivalent use of GPT or Claude outputs is prohibited.
The flow today looks something like this:
- Western frontier models → alleged unauthorized foreign extraction → Chinese open weights → lawful Western post-training
The missing direct route is:
- American frontier models → lawful Western post-training
Why does that matter? Pre-training creates a capable base model. Post-training turns it into a useful coding, reasoning, tool-using, and agentic system. A stronger teacher converts part of that expensive discovery process into a cheaper learning problem.
Distillation does not explain China’s entire open-model lead. Chinese labs have world-class researchers, substantial compute, strong pre-trained models, software-hardware codesign, and rapidly improving post-training capabilities. But distillation compresses the costly final gap between a strong base and a near-frontier system. Even if distillation represents a smaller share of a Chinese model’s total capability, it represents a meaningful share of its advantage over American open models.
New enforcement mechanisms will make large-scale distillation harder, slower, and more expensive for Chinese companies. However, enforcement will not eliminate distillation baked by state actors. Every Western frontier advance therefore creates another teacher for Chinese labs. Western builders must either reproduce those capabilities independently or wait to learn from Chinese models.
This gap gives Chinese labs a recurring structural advantage over Western companies.
The stakes extend far beyond model revenue. Suppliers of the open layer become the default base for products, synthetic data, post-training systems, evals, agents, optimization, and applied AI.
The prize is to become the substrate on which global enterprises build and improve digital intelligence.
The Weights Are Open. Our Dependency Is Not.
Downloading a Chinese model gives a Western company control over the particular version. It can run the model locally, modify it, and continue using it without permission.
But AI capability is an upgrade cycle. Western startups, model developers, and researchers increasingly rely on each new Qwen, Kimi, GLM, or DeepSeek as a stronger base, a teacher, a source of synthetic data, and a platform for further research.
If China stops releasing its strongest models, existing products will not break. They will fall behind. Reuters recently reported that Chinese authorities have discussed restricting overseas access to advanced models, including models that have not yet been released. No final policy has been announced, but Western access ultimately depends on Chinese labs and regulators continuing to publish.
There is also a more technical security problem. An open-weight model is not necessarily an auditable model.
The weights are the compressed result of training. They do not reveal the full pre-training corpus, which data was filtered or poisoned, what interventions were made during training, or whether rare trigger-dependent behavior was embedded.
We are not alleging that Qwen, Kimi, or another Chinese model contains a backdoor. The point is that possessing the weights cannot prove the absence of one. A backdoor can remain dormant during ordinary testing and activate only when an unknown trigger appears. Research shows that deliberately implanted behavior can survive supervised fine-tuning, reinforcement learning, and adversarial training.
That may be an acceptable supply-chain risk for many consumer applications. It is not acceptable for defense, intelligence, or critical infrastructure.
Open weights provide control over deployment. They do not guarantee continued access to better models, nor alignment, trust and safety in the model itself.
A Framework For A Direct American Path
American companies need a legal way to turn American frontier capability into cheaper, ownable models.
Without a domestic route, the West may lead at the closed frontier while falling into dependence on China for the open layer.
A useful framework has three parts.
- Keep building Western base models: Reflection (building open super-intelligence for enterprises & sovereigns), TML, and Nemotron are making incredible progress. But stronger pre-training alone does not solve the teacher problem. American builders also need a lawful way to absorb capabilities already developed at the American frontier.
- Create controlled teacher access: Frontier labs could sell structured training rights to qualifying Western and allied companies, whether the resulting models are released openly or deployed privately. Access could trail the frontier, cover defined capabilities, be limited to verified companies, and be metered and audited. The most sensitive biological and cyber capabilities could remain restricted. This would not allow companies to clone the newest frontier model. It would create a legal, priced route for capability transfer that sophisticated foreign actors are already pursuing covertly.
- Keep raising the cost of foreign distillation: Better identity verification, access controls, proxy disruption, and enforcement should continue. If a voluntary market does not develop, access could eventually become a condition attached to major federal AI contracts, for example.
These are starting points. Who qualifies, how far access should trail the frontier, how it should be priced, and which capabilities remain restricted are topics that deserve real debate.
Imagine if we had not allowed for training on the open web. We would have no leading AI at all. These types of policy implications are transformative. All the leading labs benefited from copious amounts of openly available data. We have to have an open and free future: It is imperative for Western competitiveness.
What should no longer go unquestioned is the current equilibrium: the West creates the frontier, part of that capability travels indirectly through Chinese models, and Western builders then depend on those models to make intelligence cheaper, adaptable, and sovereign.
To distill or not to distill is not the question. The question is whether the West creates a legal domestic path for capability transfer – or relies on an indirect path through China.