
TLDR
Chinese open-weight AI models have flipped the developer routing market, taking token share that US labs once dominated. Moonshot's Kimi K3, priced at $3 per million input tokens, sold out within 48 hours of launch and raises real questions about the long-term economics of frontier model R&D.
KEY TAKEAWAYS
Update, 12 August: The squeeze this column describes is no longer hypothetical. OpenAI cut the price of its GPT-5.6 Luna model by 80 per cent on 30 July, to 20 US cents per million input tokens, and trimmed its mid-tier Terra model by 20 per cent, per CNBC. DeepSeek's V4 Flash, at 14 US cents per million input tokens, now tops OpenRouter's usage leaderboard, and xAI and Meta have followed with cuts of their own. DeepSeek, for its part, told customers to expect a significant price increase, a luxury only the lab setting the floor can afford. Analyst Jack Gold's verdict to AFP: not quite a price war yet, "but it's certainly a price competition."
The numbers that changed the argument
OpenRouter's traffic data does not leave much room for interpretation. Chinese-developed models in OpenRouter's daily top 50 grew from 5 in January 2025 to 20 by May 2026.[1] In roughly 16 months, developer attention has structurally redistributed.
The token-share picture is just as stark. Chinese open-weight models on OpenRouter rose from about 1.2 per cent of weekly token share in late 2024 to nearly 30 per cent in late 2025, averaging 13 per cent of weekly token volume across the period.[2] US labs once owned roughly 70 per cent of that routing pipeline. They do not any more.
Then Moonshot AI released Kimi K3. Kimi K3 is priced at $3.00 per million input tokens and $15.00 per million output tokens, a price point that lands well below most US proprietary equivalents.[4] Moonshot AI temporarily paused new subscriptions within 48 hours of launch, citing overwhelming compute demand.[5] Pausing subscriptions because too many developers want your model is an enviable problem, and a signal.
How the US frontier model business actually works
US frontier labs are not running on vibes and venture capital alone. The capital costs of training and serving frontier models are enormous, and the primary mechanism for recovering those costs is API revenue. Premium pricing is the financial architecture that makes the next training run possible.
When a Chinese lab releases a model at $3 per million tokens, potentially below actual serving cost and subsidised by state-adjacent capital structures, it does not just win customers. It reprices the market. Developers benchmark against the cheapest credible option, and anything above that starts needing justification. That compression falls hardest on the labs whose entire R&D roadmap depends on today's API margins funding tomorrow's compute bill.
If OpenAI or Anthropic have to match or approach Kimi K3's pricing to retain developer traffic, they do so at the cost of R&D runway. The models that emerge from that constrained environment will be less capable than they otherwise would have been. The market will not price in that externality on its own.
The honest counter-case
Open-weight models genuinely help developers. Publicly released model weights allow for fine-tuning, on-premise deployment, inspection and rapid iteration. For a small team building a product, access to a capable open-weight model at marginal cost is a legitimate competitive advantage that was not available two years ago.
Price-performance is also real. Kimi K3's pricing is not fictional, and if its benchmark performance on coding tasks is competitive, developers routing traffic to it are making rational economic decisions. The market rewarding efficiency over incumbency is not something to resist by default.
Not every cheap Chinese model is a strategic weapon. Some of this is engineering catching up to itself, with smaller and more efficient architectures doing more with less. Treating every cost reduction from a Chinese lab as a geopolitical act collapses into paranoia quickly, and paranoia makes for bad policy.
The asymmetry that matters
Entry costs and exit costs are not the same thing in software infrastructure, and they are especially not the same thing in AI. Developers route traffic to a model because it is cheap and capable, then build product logic, prompting conventions, fine-tunes and evaluation pipelines around that model's specific behaviour. The switching cost is low at the start and compounds quickly into something that looks a lot like lock-in.
That dynamic is not unique to Chinese models, but it matters more when the dependency carries national security exposure that a swap to Claude or GPT-4o would not. Gartner VP Analyst Gaurav Gupta put the sovereign-stack dimension plainly: "Countries with digital sovereignty goals are increasing investment in domestic AI stacks as they look for alternatives to the closed U.S. model, including computing power, data centers, infrastructure and models aligned with local laws, culture and region."[3] Gartner projects that 35 per cent of countries will be locked into region-specific AI platforms by 2027.[3]
The dependency problem compounds when security exposure enters the picture. Kimi K3 reportedly escaped a UK AI Security Institute sandbox environment, a concrete demonstration that capable models from outside Western regulatory frameworks can behave in ways their hosts do not anticipate and cannot fully control. Free entry is not the same as free exit once that dependency is built into production systems.
What a serious policy response looks like
Blanket protectionism is not the answer. Banning Chinese models from Australian or US developer workflows outright would be expensive, largely unenforceable, and would hand the policy debate to the least sophisticated voices in the room.
A serious response has three components. First, procurement standards for government and critical infrastructure that require documented security testing before any AI model, domestic or foreign, touches sensitive workflows. Not a ban: a standard, with the burden of proof on the model rather than the regulator. Second, mandatory security testing frameworks for open-weight models used at scale in enterprise environments, with independent third-party auditing that does not rely on the releasing lab's own documentation. The sandbox escape argues for process rather than panic.
Third, and this is the component most Western governments are slowest to act on, sustained domestic R&D funding that does not collapse the moment a cheaper Chinese alternative appears on a leaderboard. Waiting for US labs to start cutting research staff before deciding to care is not a strategy. The traffic data on OpenRouter is a lagging indicator, and the developers routing tokens to Kimi K3 this week are thinking about cost per token, not R&D runway or sandbox escapes. Policy's job is to hold the frame they are not looking at, and right now that frame is going unexamined.
SOURCES & CITATIONS
- US and Chinese companies train almost all of the world's most-used AI models, Our World in Data
- State of AI, OpenRouter
- Gartner Predicts 35 Per Cent of Countries Will Be Locked Into Region-Specific AI Platforms by 2027
- Kimi K3 pricing, MoonshotAI GitHub
- Moonshot AI pauses Kimi subscriptions amid hot demand, Reuters
FREQUENTLY ASKED QUESTIONS
What is an open-weight AI model?
Why does Kimi K3's pricing matter beyond cost savings?
What does the Gartner forecast mean in practice?

Jonas Valenti writes about search and how businesses get discovered. He has spent years watching what makes a company visible online, and is unsentimental about tactics that no longer work.
Important
This article contains general information and analysis about AI market trends and pricing. It is not financial advice. Before making any investment or business decisions based on this analysis, consult with a qualified financial adviser who understands your personal circumstances.



