A Chinese lab shipped a model so good that demand overwhelmed its capacity — so it stopped taking new subscribers and pointed everyone at the free weights instead.
The story isn't the model.
It's that the wall was compute, and the fix was to give the model away. SCMP.
Moonshot suspended new Kimi K3 subscriptions when usage outran the GPUs it could serve.
Then it did the counterintuitive thing: announced the full open-weights drop for July 27, this Sunday. If you can't serve everyone, let them run it themselves.
"They didn't run out of intelligence. They ran out of compute."
Model capability is no longer the scarce thing — inference infrastructure is.
This is the second open-weight drop in one week (Inkling was the first). The pattern is deliberate.
Chinese labs are planting a flag on the open side — the "good-enough intelligence, low price" layer that's 90% of token volume.
And a model you've downloaded can't be switched off, rate-limited, or export-controlled by anyone.
It rhymes with the Correction: four days after the US could switch off a frontier model, an open Chinese model shipped that nobody could recall. Open weights keep showing up as the story under the story.
On July 22, OSTP director Michael Kratsios accused Moonshot on X of "large-scale distillation" — copying Anthropic's Fable to build K3.
The claim: a covert internal platform to avoid detection, running on Nvidia GB300s allegedly smuggled through Thailand, around export controls. Treasury is reportedly weighing sanctions.
"Legitimate AI distillation… plays a vital role in this open innovation ecosystem."
"However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology… is unacceptable."
Independent researchers pushed back within hours. Fable only went public July 1 — a model this strong two weeks later is hard to pin on distillation alone.
And distillation's payoff shrinks near the frontier. If it were that easy, everyone would have caught up already.
The timing is loud, either way: the accusation lands in the same week as DeepSeek V4's stable release and Kimi K3's open-weight drop — the biggest cluster of Chinese open releases yet.
Then the plot twist: the same day the White House called Kimi stolen, Nvidia's Jensen Huang went the other way — his first-ever post on X, defending it.
"Open-source models that are excellent should be used." He name-checked Kimi, DeepSeek, Alibaba and more as "world class," called the "backdoor" fear a misconception, and said the US should embrace them, not ban them. the post →
His reason is self-interested and honest: more AI use — open or not — means more Nvidia chips sold. The chip vendor and the White House now disagree in public.
Two weeks, one model, three verdicts: the White House calls Kimi K3 stolen.
Nvidia calls it excellent.
Researchers call the theft claim unproven.
The subtext is the same for all three: you can't un-ship open weights. Once they're downloaded, no accusation, ban, or export control pulls them back.
This is AI's Huawei-5G moment — provenance paranoia meets capability competition.
The US bet: export controls slow China. China's bet: open weights plus engineering velocity erase the hardware gap.
If GB300 enforcement tightens, AMD becomes the main hyperscale alternative — the chip race and the model race are now one story.
"The most capable lab of the week couldn't keep its own model online. That's not a model problem — it's an infrastructure problem, and it's everyone's."
Kimi's pause isn't a one-off. This is the week compute scarcity went mainstream — reports of Google capping Meta's API, labs rationing capacity everywhere.
Model capability stopped being the constraint.
GPUs became it.
That's the real 2026 story hiding under a model launch.