← All posts / Models

Three Open Weights Against the API: Inside Abacus.AI's Smaug Line for Enterprise Agents

Abacus.AI's new Smaug Agentic, Flash, and Mini fine-tunes of Kimi K3, DeepSeek V4 Flash, and Qwen3.8 27B beat Claude Sonnet 5 on several agentic benchmarks — with weights you can download.

Three Open Weights Against the API: Inside Abacus.AI's Smaug Line for Enterprise Agents

The most interesting argument in AI right now is not about which frontier model tops a chatbot leaderboard. It is about who gets to run the agents that enterprises actually deploy: closed APIs from Anthropic and OpenAI, or open weights that a company can download, host in its own VPC, and fine-tune for its own workloads. On September 10, Abacus.AI threw three new pieces of evidence into that argument with the launch of the Smaug line — Smaug Agentic, Smaug Flash, and Smaug Mini, a family of open-weight models fine-tuned specifically for agentic work, built on three of the strongest open-weight bases available today.

One recipe, three bases, three workloads

The Smaug line is not a new foundation architecture. It is a fine-tuning methodology applied to existing open-weight models — and that is precisely what makes it interesting. Abacus.AI’s approach combines human-curated, real-world agentic traces (logs of agents actually completing tasks) with synthetic data grounded in hard, challenging examples. The company says the recipe transfers: applied across a range of open-weight base models, it produces consistent gains on the benchmarks that matter for production agents — agentic coding, tool use, automation, and long-context reasoning.

Each model in the family occupies a different point on the capability–efficiency curve:

Smaug Agentic fine-tunes Moonshot AI’s Kimi K3, the frontier-scale multimodal mixture-of-experts base with 2.8 trillion total parameters, roughly 104 billion activated per token, 896 routed experts, and a context length of 1,048,576 tokens. This is the flagship: aimed at long coding and tool-use loops where an agent reads a repository, edits files, runs tests, interprets failures, and repeats the sequence across many dependent steps.

Smaug Flash fine-tunes DeepSeek V4 Flash 0731 and is positioned as the workhorse — the continuously running enterprise agent handling messages, documents, API calls, and workflows across many turns. Abacus.AI is candid about why it chose this base: DeepSeek Flash is efficient and robust for frequently-running agentic tasks but “susceptible to spins and confusion with long context agentic tool use.” Smaug Flash keeps the base’s one-million-token context and its serving layout, so existing serving stacks can load it without architectural changes.

Smaug Mini fine-tunes Alibaba’s Qwen3.8 27B into a compact, multimodal-capable package that fits on a single GPU — aimed at one-off bounded tasks: a customer-service agent that inspects an uploaded image, retrieves an account record, and drafts a response.

The benchmark story

The numbers Abacus.AI publishes are the real payload of the launch, because they pit open weights against the closed frontier head-on.

Smaug Flash posts a LiveBench overall score of 77.4 versus 74.2 for its DeepSeek base — and 76.0 for Claude Sonnet 5. On LiveBench agentic coding specifically, the gap widens dramatically: 61.1 for Smaug Flash, 46.8 for the base, and 59.4 for Sonnet 5. On AutomationBench (public 600, strict pass) it scores 38.83 against 34.67 for Sonnet 5; on NL2repo-bench, 73.3 against 66.34. A fine-tune of an efficient mid-scale base outscoring Anthropic’s production workhorse on agentic coding is the kind of result that makes enterprise platform teams reconsider their API bills.

Smaug Mini tells a similar story at small scale. On AutomationBench it scores 41.8 — ahead of Claude Sonnet 5 (36.2), Claude Opus 4.6 (25.5), and GPT-5.6 Luna (34.8). Its JobBench score of 50.5 (official protocol, 65 tasks) beats the base Qwen3.8-27B’s 33.4 by seventeen points, and leads Sonnet 5’s 46.4 and GPT-5.6 Luna’s 43.3.

Smaug Agentic, the largest model, “holds its own and even leads in several key benchmarks” against GPT-5.6 Sol and Claude Fable 5: GPQA Diamond 94.1 (level with GPT-5.6 Sol, ahead of Fable 5’s 92.6), AA-LCR long-context reasoning 75.7 (best in the table), DeepSWE 69.9 versus the base Kimi K3’s 67.5, and LiveBench agentic coding 64.6 against GPT-5.6 Sol’s 56.2. The closed frontier still wins Terminal-Bench 2.1 (88.8 for GPT-5.6 Sol versus 86.5) and MMMU-Pro — but the margins are thin, and the weights are downloadable.

Why fine-tuning for agents specifically

The technical premise here deserves attention. Agentic workloads differ from chat in a way that generic training does not fully address: long contexts, repeated tool calls, and the compounding problem of prompt caching breaking down across time gaps. An agent that runs for hours accumulates state, revisits earlier decisions, and must stay coherent across hundreds of steps. Base models — even very good ones — degrade in exactly these conditions, which is why DeepSeek V4 Flash “spins” in long agentic loops.

Abacus.AI’s bet is that targeted fine-tuning on real agentic trajectories closes this gap without touching the architecture. The benchmark tables support parts of that argument: the largest gains (agentic coding +14.3 points on LiveBench for Smaug Flash; JobBench +17.1 for Smaug Mini) are precisely in long-horizon, multi-step tasks. The caveat is equally clear: these are vendor-run evaluations under the company’s own harness (some reference scores are marked as “reported, not from our harness run”), and independent production evidence remains limited.

The open-weight enterprise pitch

The launch is also a strategic statement. All three models are available on Hugging Face under the licenses inherited from their base models, and Abacus.AI emphasizes that they can run inside an enterprise VPC. That reframes control over data, infrastructure, and model behavior as part of the product proposition rather than an optional compliance feature.

The important contest, in other words, is not Abacus.AI versus one model vendor. It is open-weight, self-hosted agent infrastructure against closed-model APIs that offer convenience but retain operational control. If a fine-tuned 27B model can beat Claude Opus 4.6 on automation benchmarks while fitting on a single GPU inside your own cloud, the economics of agent deployment change — and the “you must use the closed frontier” assumption weakens.

For enterprises building self-improving agents that run continuously, the Smaug line offers a credible middle path: frontier-adjacent agentic capability, open weights, and a size point for every workload. Whether the benchmark advantages survive contact with real production traffic is the question the next quarter will answer.