600 Billion Parameters, 27 Billion Awake: StepFun's Step 5 Preview Attacks the Price-Performance Frontier
StepFun's new 600B-total/27B-active MoE flagship brings a 1M-token context and $1/$2.70 per-million pricing to agentic work — with open weights promised for October 15.
On September 20, 2026, Shanghai-based StepFun (阶跃星辰) released Step 5 Preview, its next-generation flagship model aimed squarely at what the company calls the Pareto frontier of intelligence per dollar. The headline numbers: a sparse Mixture-of-Experts architecture with roughly 600 billion total parameters of which only about 27 billion are active per token, a 1 million-token context window, native multimodal input across text, images, and video, and API pricing set at $1.00 per million input tokens and $2.70 per million output tokens.
That price point deserves a beat of attention. Artificial Analysis, the independent model-evaluation service, puts the median for comparable reasoning models at $1.88 input and $10.00 output per million tokens. Step 5 Preview undercuts the output median by roughly 73 percent while, by the same service’s measure, scoring 44 on its Intelligence Index against a tier median of 24. On paper, at least, the model occupies a corner of the price-performance plane that neither OpenAI nor Anthropic currently serves.
What shipped, exactly
Step 5 Preview is available now as a hosted API on the StepFun Open Platform and through the Vercel AI Gateway, with an OpenAI-compatible endpoint, streaming, tool calling, JSON mode and JSON Schema support, prompt caching, and a selectable reasoning-effort knob (low, medium, high). A Hugging Face repository under the stepfun-ai organization exists but is empty pending the open-weights release, which the company has scheduled for October 15, 2026. A BF16 mirror posted by a third party suggests the community is already lining up: by simple arithmetic, 600B parameters means roughly 1.2 TB in BF16 before you budget a single byte of KV cache, so prospective self-hosters should be planning multi-GPU server builds now.
The target workloads are explicitly agentic: software engineering, professional knowledge work, and finance. StepFun says the model coordinated 950 web fetches within a single agent action during research tasks, and the platform documentation includes a Claude Code integration guide — a telling detail, as it positions Step 5 as a drop-in backend for the agentic tooling developers already use rather than a walled-garden assistant.
Narrow and deep, not wide and shallow
Architecturally, the interesting choice is depth. Rather than widening the network, StepFun stacked 92 Transformer layers in a narrow-deep configuration. The research team’s argument is that deeper stacks provide longer paths for implicit multi-hop reasoning — which matters most during long prefill, precisely when agents are interleaving searches, code execution, and tool returns across a million tokens of working memory.
Training leaned on on-policy, long-horizon reinforcement learning, with the company citing bit-wise alignment between training and inference across MoE routing — a nontrivial guarantee, since routing drift between train and serve is a classic source of silent quality degradation in sparse models. The engineering list also includes MTP-3 speculative decoding, FP8 quantization for the MoE layers, and KV-cache offload, together claiming more than 3x end-to-end speedup for long-horizon RL workloads. Artificial Analysis measured real-world API throughput at 99.8 tokens per second.
Benchmarks: read both columns
Company-reported numbers first. On FrontierFinance, Step 5 Preview scores 66.4 against 69.7 for Claude Opus 5 and 55 for GPT-6 Astra (with the caveat that Step 5 ran at High effort while rivals ran at Max). On DRACO, it posts 83.3 versus 87.6 and 76.8 for the same competitors. Coding tells a more sober story: 67.7 on DeepSWE v1.1, 49.0 on StepCodeBench (StepFun’s own benchmark), and 80.5 on ProgramBench — with GPT-6 Astra and Claude Opus 5 ahead on all three.
Two 24-hour agent experiments add color. In one, the model autonomously tuned an H100 kernel to 508 TFLOPS, edging out the 493 TFLOPS achieved by Claude Opus 5 under the same protocol. In the other, it took Qwen3-30B-A3B from 53.3 percent to 60 percent on AIME24 through automated post-training — a model improving another model, unsupervised, around the clock.
The independent check: Artificial Analysis scores Step 5 Preview at 44 on its Intelligence Index. That is well clear of its price tier’s median of 24, though short of the frontier models’ 60-plus club. And there is one catch the evaluation surfaced: the model generated 160 million output tokens during the index run against a 92 million median. Verbose reasoning chains eat into the per-token savings, so real-world cost depends heavily on how well you tune that reasoning-effort knob.
Why it matters
Three currents converge here. First, the agent era is shifting procurement logic from “which model is smartest” to “which model is cheap enough to run 24/7 at scale” — and a 4.5-percent active-parameter ratio is exactly the kind of economics that shift rewards. Second, Chinese labs continue to compress the gap between frontier capability and commodity pricing, and Step 5 Preview’s October 15 open-weights date will make that pressure concrete for every self-hosting shop. Third, the benchmark picture is honest enough to read: Step 5 is not beating Opus 5 or GPT-6 Astra head-to-head, but it is delivering most of the intelligence at a fraction of the cost, which is a different and arguably more commercially decisive game.
The preview label is doing real work, too. Weights are not yet downloadable, third-party evaluations are still accumulating, and the verbose-reasoning tax needs live testing against your own workloads. But as a statement of where the price-performance frontier sits in late 2026, Step 5 Preview is one of the clearest data points yet — and October 15 is circled on more than a few infrastructure calendars.
Sources
- [1] https://www.stepfun.com/step-5-preview
- [2] https://www.marktechpost.com/2026/09/20/stepfun-launches-step-5-preview/
- [3] https://aiweekly.co/alerts/stepfun-ships-step-5-preview-api-a-600b-moe-at-1270-that-scores-44-on
- [4] https://platform.stepfun.ai/docs/en/guides/models/step-5-preview
- [5] https://huggingface.co/TypeSafeAI/Step-5-Preview-BF16