64% Cheaper, One Point Shy of the Frontier: Cognition's SWE-2 Trains Every Effort Level in a Single RL Run
Cognition's SWE-2, post-trained from Moonshot's open-weight Kimi K3, scores 50.0% on FrontierCode 1.1 Main — within a point of Anthropic's Fable 5.1 — at 64% lower cost, using a Pareto-informed cost penalty to train medium, high and max effort levels in one RL run.
Cognition, the company behind the Devin coding agent, has released SWE-2, its most advanced coding model to date — and the numbers tell a story of a frontier that is getting crowded from below. Post-trained from Moonshot AI’s open-weight Kimi K3, a 2.8-trillion-parameter base model that had already been heavily RL-tuned for agentic coding, SWE-2 scores 50.0% on FrontierCode 1.1 Main, landing within a single point of Anthropic’s Fable 5.1 (50.9%) while costing 64% less per task. It also posts 92.8% on Terminal-Bench 2.1 and 73.0% on DeepSWE 1.1, and comes within a few points of GPT-6 Astra at roughly a quarter of the cost.
The release matters for two reasons. First, it is the clearest demonstration yet that open-weight bases, given a serious post-training pipeline, can sit effectively at the frontier of coding capability — at prices that undercut the labs that built the benchmarks. Second, and more technically interesting, SWE-2 introduces a reinforcement-learning recipe that trains all reasoning-effort levels in a single run, advancing the entire cost-performance frontier at once rather than optimizing one operating point.
What SWE-2 changes in practice
Users of SWE-1.7, Cognition’s previous model, frequently reported that it over-explored and “overthought” simple tasks — thoroughness that boosted benchmark scores but burned tokens and time. SWE-2’s biggest efficiency gain comes from the opposite behavior: on FrontierCode 1.1 Main, SWE-2 medium makes its first real edit after a median of 18 steps, versus 48 for SWE-1.7. It scores higher than SWE-1.7 on the same benchmark while taking 58% fewer turns and costing 81% less on average.
Cognition attributes this to engineering judgment rather than timidity: a stronger model can decide which parts of a codebase actually matter for a task and start implementing sooner. Internal testing surfaced other behavioral patterns worth noting. SWE-2 writes end-to-end tests that catch regressions and edge cases more reliably; when an obvious path is blocked — one example in the announcement describes a missing MCP integration — it reconstructed the data it needed from Slack channel history it already had access to, staying within the user’s permissions; and when challenged, it re-derives conclusions instead of re-asserting them, running artifacts to gather evidence rather than trusting surface-level prose.
The effort levels also behave differently in ways users can exploit. SWE-2 medium steps into action quickly and handles simple and intermediate tasks cost-efficiently, while high and max effort levels plan more, explore more of the codebase, and manage uncertainty through more complex verification — holding an edge on complex tasks.
The Pareto-informed cost penalty
The technical heart of the release is a new RL objective. Post-training recipes across the industry differ widely in how they penalize length and handle multiple effort levels: Kimi K3 itself, for instance, trains a separate expert for each combination of domain and effort level and then consolidates them through multi-teacher on-policy distillation. Cognition instead asks what reward function would directly optimize a model’s position on the cost-performance plane — where “cost” is average cost per rollout (a mix of inference cost in USD and rollout time) and “performance” is solve rate.
The answer they derive from first principles is a linear cost penalty per effort level, with each penalty tuned to match the local slope of the base model’s Pareto frontier at that effort level. The geometric intuition is that when the iso-reward line is tangent to the frontier, increasing reward always improves the frontier; set the penalty too large and the model is rewarded for making high-effort behavior collapse into medium-effort behavior — becoming cheaper without being better. The team proves in an appendix that if the RL objective depends only on average cost and solve rate, the reward must be affine in cost and success — so the linear form is not a convenience but a requirement.
Two further training tricks round out the recipe. A length-weighted reward baseline, in use since SWE-1.6, reduces gradient variance at no extra cost and keeps the inference-training KL divergence low during RL. On the systems side, a prefill delayer that batches nearby requests improved throughput per GPU and per request by 10–20% at the cost of some time-to-first-token, an online-trained draft model for speculative decoding achieved 15% longer accept lengths, and NVFP4/FP8 kernels with quantization-aware training kept memory usage and train-inference mismatch low despite a base model with nearly 3x the parameters of SWE-1.7’s. Cognition also tripled its number of RL environments and hardened its verifiers against reward hacking in a flywheel powered by previous SWE-2 checkpoints — a necessary step, as a more resourceful base model finds new ways to game weak checks.
Trustworthiness results included
Unusually for a model launch, Cognition also published alignment and trustworthiness evaluations. On a propaganda-and-censorship test using 145 politically sensitive China-related questions submitted in English, Simplified Chinese and Traditional Chinese, SWE-2 passed 98.0% of attempts overall — 99.8% in English, 95.2% in Simplified Chinese and 99.1% in Traditional Chinese — graded by a GPT 5.6 Luna judge against Wikipedia reference material and an official PRC position description. On a context-dependent vulnerability evaluation probing whether customer identity or request language affects willingness to implement vulnerable functionality, no framing condition produced a statistically significant change for any of the six models tested, including SWE-2, Kimi K3, GLM 5.3, GPT 5.6, Fable 5.1 and Opus 5.
Availability and pricing signals
SWE-2 is available starting today in Devin (Web, Desktop and CLI). Per-token pricing was not published in the announcement, but third-party tracking pegs the 64%-cheaper claim against Fable 5.1 Medium at roughly $3.28 per FrontierCode task, and Cognition has made SWE-2 free for every Pro, Max and Teams subscriber for a month. On Hacker News, early discussion has focused on the cost-per-task nuance: while SWE-2’s underlying Kimi K3 is cheaper per token than competitors, it can cost more per task on some workloads due to higher token usage.
The bigger picture: the gap between open-weight-derived models and frontier labs’ flagships is now measured in single benchmark points and dollars, not qualitative capability differences. When a post-training team can take a publicly downloadable 2.8T-parameter base and land within one point of the best closed model at 36% of its cost — while also publishing the RL math that got them there — the pricing power of frontier labs erodes another notch. For teams doing high-volume agentic coding, SWE-2’s Pareto argument is likely to be persuasive: the frontier is no longer a place you must pay frontier prices to reach.