Publication
The open weights debate (3D Chess)
14 min read · 27 min watch
- Capital, Markets, And Business Models
- Open Models
- Policy, Governance, And Geopolitics
Which matters more: accelerating the frontier, or accelerating access to what the frontier has already produced? The answer should be both.
Since the launch of GPT-3, AI has become an increasingly complicated game of geopolitical and commercial 3D chess.
Everyone can see the immense potential benefits of the technology. Everyone can also see the risks that come with falling behind. But the participants in this race do not all want the same thing.
The labs want better models and a lead in the race to AGI. Governments want security, power and stability. Investors want returns. Hardware and infrastructure providers want more demand and wider adoption of their technology.
When Kimi K3 launched, this game intensified. A frontier-level Chinese open-weight model does not simply add another name to a benchmark. It changes the incentives of labs, governments, investors and infrastructure providers at the same time.
Those incentives are intertwined, but they are not aligned. That is why the open-weights debate is so difficult to read—and why it needs more open discussion.
One month that changed the open-weights debate
Source1:02The debate moved quickly.
On June 10, Dario Amodei called for mandatory testing and government power to block unsafe frontier models in his essay “Policy on the AI Exponential”. Two days later, a US directive led Anthropic to remove access to Fable and Mythos.
On July 14, Demis Hassabis proposed pre-release testing and a frontier-AI standards body in “A Framework for Frontier AI”.
On July 16, Moonshot launched Kimi K3 at frontier-level performance and said the full weights would follow.
On July 17, Xi Jinping opposed single-country AI dominance and excessive national-security restrictions in remarks at the World Artificial Intelligence Conference. The official English account emphasised openness, open-source development, collaboration and sharing.
On July 22, White House science and technology adviser Michael Kratsios said on X that the United States had information that Moonshot distilled Anthropic’s Fable model while developing K3. That is an official allegation, not independent proof that the allegation is true.
On July 25, Jensen Huang made his X debut with an argument for the proliferation of open weights, supported by figures including Mark Zuckerberg and Satya Nadella.
At the time of recording, the relevant Polymarket contract put the odds of US restrictions at roughly 47%. That number was a volatile market snapshot, not a forecast to treat as fact.
This is the environment in which the open-weights debate now sits: frontier competition, national security, industrial policy and business-model protection all colliding at once.
The compute race and the $130 billion infrastructure bet
Source2:57Hassabis describes the labs as participants in an intense, multilayered commercial and geopolitical race. Competition accelerates progress and its benefits, but frontier capability is moving faster than our understanding of the technology.
That leaves the labs in a difficult position. Staying in the race requires larger compute commitments. They must forecast how much infrastructure future models will need, raise the capital, secure the compute, train and deploy the models, and then repeat the process at greater scale and on shorter timelines.
Buy too little and a lab risks falling behind. Buy too much and it risks going bust. Without this competitive dynamic, AI would not have progressed as quickly—but the financial exposure is enormous.

The compute-commitment loop turns every forecast into another, larger capital decision. Editorial illustration: The AGI Post.
During the first quarter of 2026, Amazon, Alphabet, Microsoft and Meta reported roughly $130 billion across their respective capital-expenditure measures: approximately $43.2 billion for Amazon, $35.7 billion for Alphabet, $30.9 billion of cash paid for property and equipment by Microsoft, and $19.84 billion under Meta’s company-defined measure. These figures are directionally useful, but their accounting definitions are not perfectly standardised.
Much of this spending relates to cloud, data-centre and AI infrastructure. It also sits inside a tangled network of investments, supply agreements and purchasing commitments.
Amazon has invested heavily in Anthropic, while Anthropic has made substantial commitments to AWS. Microsoft holds a major stake in OpenAI while also working with Anthropic, and Anthropic has made large commitments to Azure. Google has invested in Anthropic and supplies TPU capacity. Google also owns a stake in SpaceX, which sells compute capacity to other labs. NVIDIA invests across the ecosystem while supplying the accelerators on which most of it depends. Meta has explored leasing AI compute capacity to Anthropic.
The same companies can be investors, suppliers, customers and competitors at the same time.
Why tolerate this mess? Because the potential payout is vast. More intelligence could mean higher productivity across the economy, faster scientific discovery and better medicine. It could also mean greater cyber capability, autonomous weapons, intelligence advantages, industrial dominance and military power.
In that environment, underinvestment may look more dangerous than overinvestment—at least until the bill arrives.
Kimi K3 blows open the frontier
Source6:30Only weeks earlier, the frontier appeared close to a two-horse race between OpenAI and Anthropic. Then four major models arrived in eight days: Grok 4.5, GPT-5.6, Muse Spark 1.1 and Kimi K3.
The Artificial Analysis launch comparison showed how quickly the field had changed. Rankings and prices move, but the strategic change is harder to dismiss.
Kimi K3 appears to push the frontier across many benchmarks at a substantially lower advertised serving price. Important caveats remain. It can use more tokens on some tasks, which may erode the apparent cost advantage, and a 2.8-trillion-parameter mixture-of-experts model is impossible for almost everyone to run locally.
Even with those caveats, the significance is clear: a Chinese open-weight model is now genuinely challenging the frontier.
It is difficult to know exactly how much compute Moonshot had access to, but the company was almost certainly operating under tighter resource constraints than the leading US closed labs. A competitive model produced under those constraints simultaneously:
- pressures the economics of American closed-model providers;
- distributes Chinese technical standards and architectural choices more widely; and
- strengthens China’s position in the AI race.
That is why Kimi K3 disrupts the 3D chessboard.
Distillation, IP theft and the case for a ban
Source8:22The distillation accusations have begun, and they will continue. If the United States or its allies move to restrict Chinese open-weight models, IP theft—and distillation specifically—could become the justification.
It is worth being precise about what distillation is.
Distillation uses outputs from a stronger model to help train a weaker or newer model. It does not mean copying the model’s weights. Nor does the existence of distillation mean Chinese models are capable only because they copied American ones.
A crude distillation pipeline might work like this:
- Prepare and send millions of task prompts, such as coding requests.
- Collect the responses and, where available, the reasoning traces.
- Use that data during supervised fine-tuning to train another model.

One illustrative distillation pipeline. It explains the mechanism; it is not evidence about any specific lab. Editorial illustration: The AGI Post.
One common response is that OpenAI and Anthropic scraped the web to train their models, so it is hypocritical for them to object when competitors collect model outputs. But web-scale pre-training and collecting another model’s outputs and reasoning traces for supervised fine-tuning are not technically identical processes.

Web-scale pre-training and targeted collection of another model’s outputs are different technical processes. Editorial illustration: The AGI Post.
Anthropic’s Commercial Terms also prohibit customers from using its services to build competing products or train competing models without permission. Whether those restrictions are fair or enforceable is a separate debate. It is nevertheless easy to understand why frontier labs would try to stop potential competitors from using their outputs this way.
The more important policy question is not whether distillation happens. It is how much distillation contributes to capability.
If it is a large and growing source of Chinese model progress, closed labs and regulators will find it easier to argue for restrictions under an IP-theft rationale. If it provides only marginal gains, banning it would have limited technical effect—and using it to justify a much broader ban on Chinese open-weight models would look more like an attempt to slow a competitor.
Ben Thompson’s “Who’s Afraid of Chinese Models?” argues that the panic around Chinese open models is overstated. American frontier labs retain important advantages in capability, token efficiency, scale, distribution and the products built above the model layer.
Thompson also argues against bans on distillation. His proposed response is to allow American open-model companies to use closed frontier models as teachers, just as Chinese labs are alleged to do.
But supervised fine-tuning is only one part of the training pipeline. Distilled data may help prepare a model for later stages, but the lab still has to pre-train a base model and engineer expensive reinforcement-learning systems. Those systems require tasks, environments, trajectories, rubrics, graders and a great deal of infrastructure. Distillation does not make the rest of that work disappear.
Nathan Lambert has argued that distillation may therefore deliver only marginal gains relative to the complete pipeline. If that is right, Kimi K3 needs a broader explanation.
How Moonshot built Kimi K3
Source14:15There is probably no single magical answer.
Moonshot may have had more compute than outsiders assume, including access through different jurisdictions and a growing domestic accelerator ecosystem. Catching up can also be easier than exploring unknown territory at the frontier.
Chinese labs may have fewer product distractions. They are not all serving hundreds of millions of demanding consumer users while simultaneously pushing the frontier. There is also a growing market for training data and reinforcement-learning environments.
Then there is the simplest explanation: the teams are good.
Researchers who have spent time with Moonshot and other Chinese labs describe young, ambitious and highly capable teams. Lambert characterised Moonshot as having an exceptional culture. Bill Gurley has likewise pointed to the strength of China’s open-source culture.
Kimi K3’s own technical release supports the idea that this is not a story about one shortcut. Moonshot describes several architectural and training innovations:
- Kimi Delta Attention, intended to reduce the cost of long context;
- Attention Residuals, designed to make better use of model depth;
- extreme mixture-of-experts sparsity, allowing 2.8 trillion total parameters without activating all of them for every token;
- quantile balancing and balanced expert parallelism; and
- Per-Head Muon optimisation.

Kimi Delta Attention is one part of Moonshot’s broader whole-stack approach. Source: Kimi Linear technical report, Figure 3.
These are precisely the kinds of whole-stack innovations that become valuable under resource constraints.
Frontier performance does not come from the chip, the model architecture or the kernels in isolation. It comes from optimising the whole system together. The training pipeline increasingly resembles an advanced manufacturing process: hardware, networking, memory, kernels, data, model architecture and post-training all have to work as one system.
Kimi K3 is better understood as the output of that system than as the product of one alleged distillation attack.
Why China is going all-in on open weights
Source17:33A few days after the Kimi release, Alibaba’s Qwen team announced Qwen3.8 and said it would go open-weight. The immediately available product was a hosted Max Preview, not downloadable weights, but the direction was still significant because Qwen had kept its most capable Max models closed.

Qwen announced that Qwen3.8 would go open-weight. Source: Qwen on X.
Why would China push so aggressively toward open weights?
Part of the answer is strategic distribution. Once weights are available, each new deployment strengthens the ecosystem, spreads compatible tools and technical standards, and makes the surrounding stack more valuable. Adoption creates a flywheel.
Open weights also compound innovation. One lab’s architecture, training method or released model becomes a building block for another. Chinese labs remain fierce competitors, but they are also unusually connected. Researchers move through overlapping networks, founders mentor one another, and companies such as Alibaba invest across the ecosystem.
That structure contrasts with the closed US frontier, where leading researchers are increasingly concentrated inside organisations that cannot freely share their most important work. Closed competition protects proprietary advantage, but it also limits how knowledge compounds across the broader public.
The result is already visible: some American projects now build on, fine-tune or distil Chinese open-weight models. Thinking Machines, for example, has published work using Qwen3-8B in an on-policy distillation setup.
There is also a geopolitical layer. Xi’s framing of AI as an international public good is obviously compatible with Chinese state interests. Open weights can spread Chinese technology and influence while challenging American platform control.
But “China” is not one actor with one motivation, just as “the United States” is not. Governments, labs, investors and researchers have different incentives inside both systems. Strategic competition and genuine scientific collaboration can exist at the same time.
What the US does next
Source22:22The United States is generally supportive of open-model progress.
America’s AI Action Plan argues that openly distributed models have unique value because startups can use them flexibly without depending on a closed provider. It says America needs leading open models and should create a supportive environment for them.
The same plan also calls for evaluating Chinese frontier models for alignment with Chinese Communist Party talking points and censorship. Support for open weights and suspicion of Chinese models sit side by side.
A frontier-level Chinese open-weight model could be immensely valuable to American businesses and the rest of the world. Companies could own, modify and deploy capable intelligence without paying a closed lab for every token.
Investor Gavin Baker captured the economic tension: Chinese open models may be bad for foundation-model companies while being good for much of the wider AI economy.

The economic tension in one post: pressure on foundation-model companies, potential gains elsewhere. Source: Gavin Baker on X.
The threat to closed US labs is straightforward. Their economics depend in part on businesses continuing to rent intelligence through APIs. Those labs also need healthy margins to fund the next generation of models. Restricting Chinese competitors can therefore be framed as national security, protection of intellectual property or preservation of the domestic frontier—but it also protects incumbent economics.
The threat may still be overstated. Open weights do not eliminate the need to host and integrate models. Closed labs retain advantages in efficiency, distribution and products such as Claude Code and Codex. They can also move into specialised domains—science, biology, medicine, mathematics and engineering—where a small capability lead remains extremely valuable.
If distillation provides only marginal gains, the case for restrictions is likely to shift away from IP theft and toward cybersecurity, backdoors, censorship and national security.
Those concerns are real. Once frontier weights are released, safeguards can be removed, use is difficult to monitor and the model cannot be recalled. Under a pre-release evaluation system like the one proposed by Hassabis, open-weight models may face a higher bar because the period before release is the government’s only meaningful point of control.

Once frontier weights are released, control fragments and recall becomes difficult. Editorial illustration: The AGI Post.
NVIDIA’s “Open Weights and American AI Leadership” acknowledges that released weights are difficult to trace or reverse, but argues that “the right response to this risk is not to prohibit open weights.”
Dean Ball offered a more pessimistic prediction: the US may create regulatory risk around the use of Chinese open-weight models.
The incentives remain mixed.
Governments gain more control when frontier labs stay closed. Closed labs protect the margins that fund their next generation of models. On the other side, the rest of the economy gains access to cheaper, more ownable intelligence. Open models can spread capability through the long tail of businesses and users that will never train a frontier model themselves.
Which matters more: accelerating the frontier, or accelerating access to what the frontier has already produced?
The answer should be both.
The difficult part is managing the dance between them.
Sources
Section 08- Dario Amodei — Policy on the AI Exponential
- Anthropic — Fable and Mythos access notice
- Demis Hassabis — A Framework for Frontier AI
- Moonshot AI — Kimi K3 launch
- Chinese Ministry of Foreign Affairs — Xi Jinping’s AI-governance remarks
- Michael Kratsios — Moonshot distillation allegation
- Artificial Analysis — Four frontier launches in eight days
- Anthropic Commercial Terms
- Ben Thompson — Who’s Afraid of Chinese Models?
- Thinking Machines — On-policy distillation
- America’s AI Action Plan
- NVIDIA — Open Weights and American AI Leadership
- Gavin Baker — Kimi K3 and the broader AI economy
- Dean Ball — prediction on regulatory risk
Tags
- AI Geopolitics
- AI Policy And Governance
- Frontier Lab Business Models
- Open Weights