← All posts / Models

Same Answers, Half the Tokens: How Fireworks Trained Ember-1 to Stop Overthinking

Fireworks Research rebuilt Kimi K3 into Ember-1, a specialized model that cuts reasoning tokens by up to 71% while matching or beating the original on coding and agent benchmarks — and it dominated Hacker News this weekend.

Same Answers, Half the Tokens: How Fireworks Trained Ember-1 to Stop Overthinking

The most expensive part of a modern reasoning model is not the answer — it is the thinking. Frontier reasoning models like Kimi K3 routinely spend more than 90% of their generated tokens on internal reasoning traces rather than the final response, and in multi-turn agentic workloads the bill compounds brutally: every new turn replays all prior reasoning back through the model, so context — and cost — grows roughly quadratically with the number of turns.

This week, Fireworks AI launched Ember-1, a specialized model from its new Fireworks Research division that takes a direct swing at that economics problem. Built on top of Kimi K3, Ember-1 was trained to reason more efficiently: it produces roughly 40% fewer tokens overall — with reasoning-token reductions reaching 71.3% on agentic tasks — while holding or improving answer quality across seven public benchmarks and two customers’ live production traffic. The launch hit the front page of Hacker News over the weekend, gathering more than 500 points and 220 comments, largely because it touches a nerve every AI engineer currently feels: inference cost at agentic scale.

The problem: thinking models think too much

Fireworks’ starting observation is one any team running agentic coding workloads will recognize. Kimi K3 is one of the strongest open reasoning models available, but its long thinking traces made automated coding expensive at scale. The obvious fix — turning down K3’s reasoning-effort setting — didn’t work: lower effort gave up too much quality. The model didn’t just need to think less; it needed to learn to think efficiently, which meant retraining it.

Crucially, not all reasoning is waste. Self-reflection — revisiting an assumption, responding to environment feedback, tracing an outcome back to an earlier decision — is how a model recovers from mistakes. The Fireworks team’s insight was that the excess reasoning beyond that useful core can be removed without touching the final answer, and that this behavior can be taught rather than dialled back with a knob. Their experiments showed K3’s reasoning could be shortened by 35–50% with no accuracy sacrifice.

Fifty experiments, two hundred evaluations

Getting there was not a prompt tweak. Fireworks Research ran more than 50 training experiments and over 200 evaluations, developing new training algorithms along the way to shorten reasoning without losing accuracy. The training collection spanned mathematics, coding, instruction following, conversation, search, tool use, and software engineering — covering both standalone problems and extended multi-turn interactions, with task feedback guiding on-policy planning so the model learns to adapt its reasoning to observations and outcomes.

One detail worth noting for the enterprise crowd: Fireworks states it used its own data and no customer data to train the model. The entire effort ran on Fireworks Serverless Training, which the company credits for the fast iteration — no GPU provisioning, pay-per-run, and a research-to-launch cycle in a fraction of the usual time and cost.

The numbers: on or near the Pareto frontier everywhere

On the coding side, Ember-1’s headline results against Kimi K3 at various reasoning efforts:

BenchmarkNK3 LowK3 HighK3 MaxEmber-1Cost vs K3 Max
Terminal Bench 2.18976.4%77.6%80.9%82.0%−51.9% / −$23.1
SWE-bench Verified50080.4%86.0%93.2%92.2%−15.5% / −$68.1
SWE-Interact756.7%13.3%21.3%20.0%−32.5% / −$60.8
DeepSWE 1.111355.8%62.8%66.4%75.2%−23.7% / −$126.9
τ-2 Bench Airline5064%64%64%66%−5.9% / −$0.3

Read that carefully: on Terminal Bench and DeepSWE, Ember-1 doesn’t just match K3 at max reasoning effort — it beats it, at roughly half to a quarter of the cost per task. The cost figures were computed using public Kimi K3 API pricing ($3/M uncached input, $0.30/M cached input, $15/M output). Across every benchmark with more than 50 samples, Ember-1 sits on or near the cost-quality Pareto frontier, matching K3-max quality at a fraction of the spend and strictly dominating K3-low.

Fireworks also evaluated Ember-1 on Doximity’s Bedside Bench, a physician-validated benchmark of 500 clinical cases across 10 categories, as part of its newly introduced Specialized Intelligence Index (SII). There, Ember-1 set a new Pareto frontier across both open and closed models — including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 — on cost per task. The company’s broader analysis of GPT-6 Astra, Claude Opus-5, and GLM 5.3 found Ember-1 leading on that frontier.

Real traffic, not just benchmarks

The most convincing validation came from live A/B tests on two customers’ production coding workloads: approximately 35% fewer tokens per task at comparable quality, with downstream product metrics — task completion, success scores, failure rates — holding or improving. One of those customers has since moved Ember-1 into live production and plans to scale it to fully replace the base model.

On Fireworks’ own internal developer traffic, the agentic numbers were even starker:

ModelScoreStepsOutput tokensReasoning-token cutTotal-token cut
Kimi K30.75123.849.3K——
Ember-10.75321.429.9K71.3%39%

Same score, fewer steps, 40% fewer output tokens — and 71% of the reasoning overhead gone. Fireworks’ proudest internal result, they say, was “no news”: developers didn’t notice the model had been swapped while their token consumption quietly dropped. For a product whose value proposition is “same answers, fewer tokens,” an invisible rollout is the strongest signal you can get.

Research Preview, and a signal about where open models go

Ember-1 is rolling out as a Research Preview on Fireworks Serverless, serving alongside base Kimi K3. The company is also introducing a release model where new research models get two-week serverless access and become permanent based on community demand — a notably demand-driven approach. Fireworks is additionally launching training support for Ember-1, letting enterprises fine-tune token-efficient variants on their own workloads.

The Hacker News discussion surfaced the bigger question: can these techniques generalize? Commenters immediately asked for the same treatment applied to DeepSeek’s Flash models and Qwen 3.8, which suffers similar overthinking behavior at the small-model end. If efficient-reasoning training becomes a commodity layer on top of open base models — the way quantization and distillation did — the open ecosystem gains a compounding economic advantage that closed labs, which monetize raw token volume, have less incentive to pursue.

That framing is what makes Ember-1 more than an incremental launch. The frontier race has mostly been about capability; Ember-1 argues the next competitive axis for open models is cost per correct answer. For anyone running agentic workloads where reasoning tokens dominate the bill, that is the metric that actually matters.