← All posts / Industry

Kimi Splits in Two: Moonshot Separates General and Coding Memberships as Demand Overwhelms Compute

Moonshot AI is splitting Kimi subscriptions into separate general and coding tiers to ration scarce GPU capacity — the clearest sign yet that open-weight success has a compute bill attached.

Kimi Splits in Two: Moonshot Separates General and Coding Memberships as Demand Overwhelms Compute

Moonshot AI is about to make its most consequential product decision since launching Kimi K3: splitting the Kimi subscription into two separate memberships, one for general use and one for coding. A banner on kimi.com has warned since August 20, 2026 that “new membership tiers are coming” that separate everyday chat access from coding-focused access, while promising that current subscribers will not be affected.

On the surface this reads like a routine pricing change. It is not. It is the latest and clearest move in a pattern that has defined Chinese AI labs throughout 2026: ship a headline-grabbing open model, then scramble to manage the compute required to serve everyone who wants it.

What is actually changing

Under the current system, a single Kimi membership bundles everything: chat on web and mobile, Kimi Work document features, and access to Kimi Code, the company’s coding agent service built on the K3 model. The new structure splits those into two products:

  • Kimi Membership — covering Kimi Web, the mobile app, and Kimi Work
  • Kimi Code Membership — covering the coding agent workflows, including Kimi Code CLI usage and API benefits on third-party platforms

Moonshot confirmed the direction in a July statement on X: “Going forward, we’ll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership” for coding workflows. The company’s help center now carries a notice that “a new membership system is coming soon. Kimi membership benefits and Kimi Code benefits will be split so you can purchase each separately, as needed.”

The exact prices of the new tiers have not been finalized publicly. Kimi’s existing paid plans run roughly $19 to $199 per month, with annual billing bringing the effective range down to about $15 to $159. A briefly-published pricing page that leaked the split before being replaced with a “coming soon” notice showed separate tiers, but Moonshot swapped it out before any prices could be confirmed — a sign the company is still calibrating.

Why Moonshot is doing this: the compute squeeze

The split did not come from a product roadmap. It came from a capacity crisis.

Kimi K3, launched in mid-July 2026, is a 2.8 trillion parameter open-weight model — the largest open-weight AI system released to date. Demand outran every forecast. Within roughly 48 hours of launch, GPU utilization pushed against the ceiling, and on July 19 Moonshot paused new Kimi subscriptions entirely, saying K3 “has received far more love than we expected.” The pause notice drew millions of views.

That subscription pause was the emergency brake. The membership split is the permanent fix. Coding workloads — agentic terminal sessions, long multi-file refactors, continuous CLI usage — are vastly more compute-intensive per user than chat queries. By unbundling coding from general access, Moonshot can price and ration the two demand curves independently, the same way an airline separates carry-on from checked baggage: the person bringing more weight pays for it separately.

A help-center note from July makes the constraint explicit: “Due to computing capacity constraints, the new plans are not yet available. The exact launch time is still to be determined.”

The K3 context: why demand exploded

The split only makes sense against the scale of Kimi K3’s success. The model holds a 1 million token context window and sits at the top of open-weight leaderboards, which made it an instant default for developers who want frontier-class capability without frontier-class API pricing. Kimi Code — the CLI and coding agent product built on it — supports using membership benefits in both the official client and third-party platforms, and K3 shipped in Kimi Code on day one.

The GPU math is unforgiving. A 2.8 trillion parameter model serving agentic coding sessions consumes orders of magnitude more inference compute than a chatbot answering questions. When your model is open-weight and your subscription base is growing faster than your data center capacity, something has to give. Moonshot chose to unbundle rather than degrade everyone’s experience equally.

The IPO backdrop

There is also a financial storyline running underneath. Moonshot is preparing a Hong Kong listing, with reports pointing to a filing as early as the end of 2026 and a pre-IPO round targeting roughly $50 billion — up from a $30 billion valuation discussed just a month earlier, according to SCMP. Bloomberg reported in July that the company told investors it was seeking their blessing for a listing within six months. (A later TechNode report, citing an insider, said an August filing rumor was untrue — timing remains fluid.)

For a company walking into an IPO, the membership split does double duty. It solves an engineering problem — rationing scarce inference compute — while also making the revenue story cleaner. Two separately-purchasable memberships mean two ARPU lines, clearer cohort economics, and a coding product that can be valued as its own business rather than a feature buried in a general subscription. Investors reading a prospectus will find that considerably easier to underwrite.

What it means for users and the industry

For existing subscribers, nothing changes — Moonshot has been explicit on that point. For new users, the practical effect is that a developer who only wants Kimi Code no longer has to pay for the general product’s features they don’t use, and vice versa. Whether that lands as a win depends entirely on where Moonshot sets the prices; the community backlash in July, when a draft plan briefly appeared and was pulled back, suggests the company knows it is walking a sensitive line.

The broader signal matters more than the pricing details. The prevailing narrative of 2026 has been that open-weight models are commoditizing frontier AI, with Chinese labs giving away what Western labs charge premium prices for. The Kimi split is the counterweight: open weights don’t make inference free. Serving a frontier-class model at scale still requires gigawatt-adjacent infrastructure, and when demand outstrips supply, even the most open-friendly lab has to ration.

Moonshot is not alone. DeepSeek moved V4-Pro to general availability in August while introducing peak and off-peak billing that raised some output token prices sharply — a company long known for near-free pricing starting to charge closer to what serving a frontier model actually costs. Z.ai is gating GLM-5.3’s open weights behind a staged safety review. Across the board, the free-and-open era of frontier inference is giving way to something more structured: capacity windows, tiered access, and product lines that separate the casual user from the compute-hungry one.

What to watch

The open question is execution. The July draft-plan backlash, the subscription pause, and the “still to be determined” launch timing all point to a company managing expectations carefully. Watch three things: the final price points when the new tiers actually go live, whether Kimi Code membership includes meaningful API or third-party platform benefits that justify its cost, and how the split interacts with the Hong Kong IPO timeline.

One thing is certain: the era of one-subscription-fits-all for frontier open-weight models is ending. Moonshot is simply the first major lab to make that explicit.