← All posts / Models

A 600B Model With 27B Active: StepFun's Step 5 Preview Matches Kimi K3 at a Seventh the Price

StepFun skips Step 4 entirely and drops Step 5 Preview: 600B sparse MoE, 27B active per token, 1M-token context, AA Intelligence Index 44 level with Kimi K3 — at $1/$2.70 per million tokens. Weights land October 15.

A 600B Model With 27B Active: StepFun's Step 5 Preview Matches Kimi K3 at a Seventh the Price

The most interesting number in Chinese AI this week is a fraction: 4.5 percent. On September 20, Shanghai-based StepFun (阶跃星辰) opened API access to Step 5 Preview, a 600-billion-parameter sparse mixture-of-experts flagship that activates only 27B parameters per token — an unusually sparse ratio even by MoE standards — and pairs it with a 1-million-token context window and native image input. The company skipped its Step 4.x line entirely, jumping from Step-3.7-Flash straight to Step 5, which is itself a statement about how large a step it believes this to be.

The result, at least on the Artificial Analysis Intelligence Index, is a model that lands at 44 — level with Moonshot’s 2.8-trillion-parameter Kimi K3, one point behind GLM-5.3 and Claude Opus 5 (both at 45), and well ahead of DeepSeek V4.1 Flash (40) and V4 Pro (36). That is frontier-adjacent performance from a model whose active parameter count puts its per-token serving cost in a much smaller model’s band. The pricing follows the same logic: $1.00 per million input tokens and $2.70 per million output tokens, against medians of $1.88 and $10.00 across comparable models, with a 95% cache discount. StepFun frames Step 5 Preview as roughly a seventh the cost of GPT-5.6 Sol at equal index performance.

What shipped, and what didn’t

A precise reading of this launch requires separating two dates. The API and StepFun Studio opened on announcement day, September 20; the weights did not. The Hugging Face repository stepfun-ai/Step-5-Preview-BF16 existed by the afternoon of the announcement and contains exactly one file — a .gitattributes — with no weights, no license, no model card and an empty config object. The Hugging Face API record shows a creation timestamp of 2026-09-20T03:52Z, zero downloads, and no sibling repositories for FP8, NVFP4 or GGUF variants of the kind that accompanied Step-3.7-Flash in June.

StepFun’s own materials put the open-weights release on October 15, 2026. The BF16 suffix in the repository name is the tell: a lab that intends to publish a bfloat16 checkpoint first — the format you fine-tune and quantise from, before the serving builds arrive — reserves the namespace at creation time and fills it later. Until then the repository proves exactly one thing: that StepFun has claimed the name. Any deployment plan that assumes a self-hosted Step 5 Preview before mid-October is a plan with no artifact behind it.

The numbers, sorted by who ran them

Two classes of figure circulate around this launch, and they should not be read the same way.

Independent (Artificial Analysis, measured): Intelligence Index v4.3.2 score of 44, ranking 24th of 200 models scored. Terminal-Bench 4.0 at 33.3% — ahead of GLM-5.3 Flash, close to GPT-5.6 Terra, and far ahead of both DeepSeek V4.1 Flash (26.8%) and Kimi K3 (12.6%). Output speed of 99.8 tokens per second (above the ~70 average) and time to first token of 2.96 seconds. On the AITIER coding board, Step 5 reportedly edges past Claude Fable 5.1.

Vendor-stated (StepFun’s own harnesses): a single-task cost claim of one-eighth of Claude Opus 5 on ALE-CLI, FrontierFinance and DRACO agentic benchmarks; a 24-hour GPU kernel optimisation run that lifted an MLA kernel to 508 TFLOPS peak against Claude Opus 5’s 493; an automated post-training experiment that moved Qwen3-30B-A3B from 53.3% to 60% on AIME24; and DeepSWE v1.1 at 67.7% on its in-house StepCodeBench (49.0%) — a benchmark StepFun built itself from 553 repositories, nine task types and 33 programming languages, which makes it a reasonable instrument for measuring its own model and not an independent one.

The pattern from this year’s Chinese-lab flagships is consistent: the aggregate third-party index holds up, and the vendor’s headline rows drift. The aggregate is the number to plan against.

The verbosity tax

One independent number cuts against the price tag and deserves equal billing. The Artificial Analysis index run consumed 160 million output tokens for Step 5 Preview against a median of 92M across models scored — the model is very verbose. At $2.70 per million output tokens, that verbosity is not free: the output tokens alone account for roughly $432 of the $922.84 total evaluation cost. A low headline price and a high token count per task are two halves of one number, and the figure that already contains both — $0.71 per Intelligence Index task — is the one to compare across models. On that blended metric Step 5 Preview remains well priced, but the gap to cheaper models narrows once their terseness is accounted for.

Positioning: agentic work, not chat

StepFun’s framing is narrow and deliberate: Step 5 Preview is a base model for real-world agentic work — AI coding, software engineering, professional knowledge work and finance — with the emphasis on long context, multi-turn tool calls and sustained execution rather than chat quality. The 1M-token window and the sustained-execution benchmarks are the substance of that claim; the chat experience is not the pitch.

One unresolved contradiction matters for anyone building on it early: a separate third-party configuration circulating in developer tooling lists the model at 350,000 tokens of context with a 64,000-token maximum output, while StepFun’s announcement says 1M and mentions no output cap. A 64k output ceiling would be a hard constraint on exactly the long-horizon agentic work this model is being sold for. That is a question for the vendor’s documentation, not for a leaderboard.

Why the sparse ratio is the story

The 600B/27B split — roughly 4.5% active — is the mechanism behind everything else in this launch. At 27B active, per-token compute cost sits in the band of a mid-sized model while the 600B total remains mostly a memory-footprint question. StepFun presents this as three generations of a single pursuit of the compute-capability Pareto frontier: Step-3.5-Flash (196B total/11B active), Step-3.7-Flash (198B/11B), and now Step 5 at 600B/27B. It is a coherent story, and the independent index scores give it more credit than marketing usually earns.

For Chinese developers, the domestic Token Plan sweetens it further: entry at ¥49 for what amounts to ¥400 of quota, with no five-hour rate limit — terms that compare favorably with international subscription plans.

What to watch before October 15

Four checkable things decide whether Step 5 Preview becomes a model you self-host, one you rent, or one you benchmark once and move past. Does the repository fill — and under what license? Does a published config confirm 600B total and 27B active? Does StepFun state the output ceiling alongside the context window, resolving the 350k/64k question? And does the price hold once the preview suffix comes off?

The precedent is favorable: StepFun’s Step-3.5-Flash and Step-3.7-Flash both shipped as genuinely open checkpoints with paper-grade documentation, and the company’s GOT-OCR 2.0 built real community goodwill. If October 15 delivers a BF16 checkpoint under a permissive license with the specs as announced, a 44-index model at $1/$2.70 — or free on your own hardware — becomes one of the strongest efficiency plays in the open-weights ecosystem. Until that repository holds more than a .gitattributes file, the announcement is the beginning of the story rather than the end of it.