← All posts / Models

Moonshot Retires Kimi K2.5, Bets the Company on K3 at 5x the Price

Moonshot AI completed the retirement of Kimi K2.5 and moonshot-v1 on August 31, leaving the 2.8T-parameter Kimi K3 as its only flagship — at roughly 5x the per-token price of its predecessor, a deliberate reversal of China's year-long price war.

Moonshot Retires Kimi K2.5, Bets the Company on K3 at 5x the Price

Moonshot Retires Kimi K2.5, Bets the Company on K3 at 5x the Price

Somewhere in a Beijing control room on Sunday night, someone at Moonshot AI pressed the button that China’s AI industry spent two years refusing to press. On August 31, 2026, the API for Kimi K2.5 — the trillion-parameter multimodal flagship the lab released and open-sourced in January — went dark, taking the older moonshot-v1 series with it. As of Tuesday morning, Moonshot’s product catalog contains exactly one current model: Kimi K3, the 2.8-trillion-parameter heavyweight open-sourced in July. And K3 costs roughly five times more per token than the model it just replaced.

On the surface, this is routine model lifecycle management — every lab from OpenAI to Google has sunset old checkpoints. What makes Moonshot’s move remarkable is the pricing arithmetic stacked on top of it. Kimi K2.5 was served at around $0.60 per million input tokens and $3.00 per million output tokens on Artificial Analysis’s tracking. Kimi K3 lists at $3.00 per million input and $15.00 per million output — a 5x jump on both axes, making it, as several outlets noted at launch, the most expensive release ever from a Chinese AI lab. Retiring the cheap models the same month the pricier flagship becomes the only option is not a transition. It is a statement.

The end of the price-war chapter

For context on just how sharp a reversal this is: throughout 2025 and deep into 2026, Chinese labs competed almost exclusively on rock-bottom pricing. DeepSeek’s aggressive per-token cuts set the reference point; Qwen, GLM, and Kimi followed, each release landing cheaper than the last while benchmarks kept climbing. The strategy made sense for market entry — undercut Western frontier pricing by an order of magnitude, capture developer mindshare, and treat compute subsidies as customer acquisition cost. Kimi K3’s July launch already cracked that consensus, arriving at prices that matched Anthropic’s Sonnet-class tiers rather than undercutting them. But as long as K2.5 remained available at a fifth of the cost, K3’s pricing looked like an experiment — one option in a menu.

Now the menu has one item. Moonshot is forcing every developer who built on the cheap tiers to make a binary choice: pay frontier-adjacent prices for K3, or migrate to a competitor’s model. According to Moonshot’s own announcement thread on Weibo, new users were already locked out of kimi-k2.5 and moonshot-v1 as of mid-July; the August 31 deadline merely completed the shutdown for existing workloads. Tencent Cloud, which hosted K2.5 for its own customers, took the model offline at 00:00 Beijing time on August 31 and directed users toward Kimi K2.6 — a transitional iteration that, notably, Moonshot itself is not presenting as the flagship path forward.

What developers are actually being asked to buy

The case for K3 is not hand-waving. The model was the largest open-weight release in history when it landed — 2.8 trillion parameters built on the KDA hybrid linear attention mechanism, with native vision understanding and a 1-million-token context window. It generated the largest open-source shockwave of the summer: Fortune reported that K3’s benchmark performance (“competitive” with Anthropic’s Fable 5, “substantially outperforming” Opus-class models on several suites) helped trigger an estimated multi-trillion-dollar swing in semiconductor market value the week it shipped. On Artificial Analysis’s independent tracking, K3 sits at the top of the open-weight leaderboard with an Intelligence Index around 60, ranks second on WebDev Arena, and third on the Agent leaderboard. Moonshot’s K3 achieved a 13-point Intelligence Index jump over K2.6 — the model generation in between.

The catch, as always, is in the fine print of self-hosting. K3’s open weights ship as 96 separate files totaling roughly 1.56 terabytes under a custom (Modified MIT) license. That puts genuine self-hosting out of reach for all but the largest operators, which means that for most developers “using Kimi” in practice means paying Moonshot’s first-party API rates — the same rates that just went up 5x. The open-weight label and the pricing power are, in this design, fully compatible: the weights being public costs Moonshot little when almost nobody can serve them cheaper than Moonshot can.

Why now

Three forces plausibly converged on this decision. First, unit economics. Reporting throughout the summer — including The Information’s coverage of DeepSeek’s approaching $7.4 billion raise and the general financing environment — has made clear that subsidized inference at scale is a luxury only labs with fresh capital can sustain indefinitely. Moonshot reportedly explored its own fundraise this year; a catalog that monetizes at frontier rates is a stronger story to investors than one that monetizes at a fifth of them.

Second, the prize moved. The battleground in late 2026 is not cheap chat API calls — it is agentic workloads: multi-step coding, browsing, tool use, tasks that burn hundreds of thousands of tokens per job. K3’s profile (1M context, third place on the Agent leaderboard, swarm-style sub-agent coordination inherited from the K2.5 lineage) is aimed squarely at that market, where per-job value is high enough that $15 per million output tokens still undercuts Western frontier rates on a per-task basis. Moonshot is not trying to win the race to the bottom anymore; it is trying to win the agent workload, where the money actually is.

Third, competitive positioning. With DeepSeek, Qwen, GLM, and Tencent’s Hy series all shipping strong open-weight models this summer, differentiation by price alone had become a commodity. Every one of those labs can match a price cut. None of them can easily match K3’s specific combination of scale, context length, and agent-suite rankings. Retiring the budget tiers converts Moonshot’s story from “cheapest capable model” to “the open-weight frontier” — a category of one, for now.

The risks of a one-model catalog

The bet has obvious failure modes. Developers who chose Kimi for cost reasons — a population built over two years of price war — have no native upgrade path, and switching costs in the API world are asymmetric: migrating to Qwen, GLM, or DeepSeek is a config change, not a rewrite. If K3’s agentic performance advantage erodes as rivals ship their next iterations (Qwen’s architecture preview this week suggests Qwen4 is close), Moonshot will have burned its budget-tier goodwill for a lead that may last a quarter or two.

There is also the trust dimension. An eight-month service life for a flagship model — January release to August retirement — sets an aggressive deprecation cadence that enterprise buyers notice. Building a production system on a model that may vanish in under a year requires either strong migration tooling or contractual assurances, and Moonshot has not traditionally been the player enterprises look to for either.

And the strategy implicitly concedes the volume game. Alibaba’s Qwen, by contrast, runs a full-spectrum strategy — from 2.4T-parameter Max down to sub-10B edge models — which is precisely why it dominates open-weight download counts while Moonshot’s pure-frontier portfolio pulls a fraction of the volume. One model means one point of failure, one price point, and one bet on where the market’s center of gravity lands.

What to watch

The verdict on this decision will arrive quickly, and there are three concrete signals to track. First-party API traffic: if K3’s usage share on trackers like OpenRouter holds or climbs through September, the price increase stuck. The K2.6 question: whether Moonshot keeps the cheaper transitional model alive on partner clouds like Tencent’s as a de facto budget tier, quietly softening the one-model narrative. And the next Qwen/DeepSeek/GLM release: if a rival matches K3’s agent rankings at half the price within the quarter, Moonshot’s flagship-only posture will look like it peaked on day one.

What is already clear is that a phase of the industry has formally ended. The era in which Chinese labs could be counted on to undercut everyone on price — the assumption underlying a year of Western pricing anxiety and at least one front-page market shock — concluded not with a price hike announcement, but with the quiet deletion of the cheap models themselves. Moonshot has decided that it would rather be scarce and expensive than ubiquitous and subsidized. The rest of the Chinese labs now get to show whether they agree.