IBM's Granite 4.2 Brings Open-Weight Reasoning to Enterprise Agents
IBM's new 3B/8B/30B open-weight models add native thinking modes and agentic RL training, aiming reasoning at on-prem enterprise workloads.
IBM has released Granite 4.2, a family of dense reasoning language models available in 3B, 8B, and 30B parameter sizes, purpose-built for the agentic workflows that enterprises increasingly demand. Released on August 25, 2026 under an Apache 2.0 license, the models arrive at a moment when the AI market is splitting into two distinct tracks: frontier labs racing toward ever-larger models, and enterprises that want smaller, controllable systems running on their own infrastructure. Granite 4.2 is IBM’s clearest statement yet about which side of that divide it intends to win.
What’s in the release
The Granite 4.2 family consists of three language models — 3B, 8B, and 30B parameters — all dense architectures rather than mixture-of-experts. Each model supports a context window extended to 512K tokens through a five-phase training schedule, and each was pre-trained from scratch on roughly 15 trillion tokens.
The headline feature is native reasoning. Granite 4.2 models include “thinking” capabilities — step-by-step reasoning that helps them plan before acting, weigh trade-offs before deciding on a path, and catch mistakes before those mistakes play out in real systems. The thinking mode can be toggled on or off, giving developers a latency-versus-accuracy dial for each workload. There is also a low-effort reasoning mode that sits between the two extremes.
For the 8B and 30B models, training went further with a specialized “agentic RL” phase focused on enterprise-style tasks: software engineering, terminal-based coding, and search-driven workflows. The result, according to IBM, is models that can navigate codebases, handle multi-step development tasks, and operate in terminal environments — the daily reality of how software actually gets written and shipped inside companies.
The training recipe behind the numbers
IBM attributes the capabilities not to scale alone but to a redesigned training process. Building on the Granite 4.0 foundation models, the team introduced an expanded, multi-stage reinforcement learning regimen.
Training begins with supervised fine-tuning, then progresses through several RL phases. The first stage, “foundational RL,” was applied across all three model sizes and strengthens mathematics, science, coding, reasoning, and tool calling. It combines verifiable rewards with reward-model-based evaluation, so models learn both correctness and higher-level quality signals. The 8B and 30B models then continue into the agentic RL phase, followed by reinforcement learning from human feedback alignment.
Two further innovations shaped the coding and reasoning improvements. The models were trained on 1 trillion tokens of synthetic code generated through IBM’s CodeAlchemy pipeline, and put through an intermediate “mid-training” step shown to unlock additional reasoning capability. On the serving side, a speculative decoding layer lets the models output text faster while serving more concurrent users — a direct attack on inference operating costs.
On benchmarks, the 30B model scores 57.00 on SWE-Bench Verified, with the 8B at 47.67 — respectable numbers for models small enough to run without a data-center-scale GPU fleet. IBM is also working with Hirundo to apply machine unlearning post-training, targeting and reducing undesirable outputs without full retraining.
Why enterprises want small reasoning models
The strategic logic of Granite 4.2 becomes clear against the backdrop of 2026’s AI economics. Frontier model APIs are powerful but carry unpredictable costs, data-residency questions, and vendor dependence. Meanwhile, agentic workloads — where a model might make hundreds of tool calls across a long session — multiply token consumption dramatically.
A 3B or 8B model with genuine reasoning ability changes that math. High-throughput agentic tasks can run on the smaller models efficiently, while the 30B is reserved for deeper reasoning and complex coding workflows. Because the weights are Apache 2.0, organizations can download, fine-tune, and deploy them in production without licensing restrictions, across cloud, on-premises, and edge environments.
The models ship with native tool calling, coding support, and are distributed through Hugging Face, Ollama, GitHub, LM Studio, and major inference providers — meeting enterprises and hobbyists alike where they already work.
Speech models ride along
The release also includes two new speech models that represent a structural break from the previous generation: Granite Speech 5.0 Turbo CTC and 5.0 Turbo CTC NC. At just 470 million parameters, they are among the smallest models in the Granite family and target laptops, smartphones, and edge devices.
Unlike earlier Granite Speech models, they have no LLM backbone. Using connectionist temporal classification (CTC), they map audio to text efficiently and learn directly from raw audio and text pairs. The speed results are striking: current leaders on the Hugging Face Open ASR leaderboard process at around 6,000 RTFx, while IBM measured Granite Speech 5.0 Turbo CTC at roughly 12,600 — on a single H200 GPU. Reduced sampling means the model can transcribe three hours of recorded voice in about a second, making it practical for high-volume call-center analytics and real-time transcription running locally on a laptop. A non-commercial variant trained on restricted-use data is also available.
The open-weight reasoning race
Granite 4.2 lands in an increasingly crowded field of open-weight reasoning models, but IBM’s positioning is distinct. Rather than chasing leaderboard supremacy against trillion-parameter frontier systems, the company is optimizing for the constraints that actually govern enterprise adoption: model size, deployment flexibility, licensing clarity, and inference cost.
That bet reflects a growing consensus that the next phase of AI value creation happens not in the demo but in production — where agents must call the right tools in the right order, verify their own output, and do so at a unit cost that survives contact with a real budget. Whether Granite 4.2 becomes the default backbone of enterprise agents remains to be seen, but it marks a credible open-source answer to the question many CTOs are now asking: can we get reasoning without renting it forever?
For teams evaluating the release, the models are available today on Hugging Face, Ollama, GitHub, and through watsonx.