← All posts / Models

Beam Lands: Reflection AI Ships a 501B-Parameter Open-Weight Model That Matches GLM-5.2 on a Fraction of the Compute

Reflection AI has officially unveiled Beam, a 501B-parameter sparse MoE model with 23B active per token, SWE-bench Verified at 80.9, and reasoning on par with Z.ai's GLM-5.2 using 3-4x less inference compute — the first credible American answer to China's open-weight dominance.

Beam Lands: Reflection AI Ships a 501B-Parameter Open-Weight Model That Matches GLM-5.2 on a Fraction of the Compute

For two years, the open-weight AI race has been a one-sided story: the most powerful freely downloadable models — DeepSeek, Qwen, GLM, Kimi — came almost exclusively from Chinese labs, while America’s frontier players kept their best systems behind API toll booths. On October 5, 2026, that changed. Reflection AI, the Nvidia-backed startup founded by ex-DeepMind researchers, officially unveiled Beam — its first open-weight model, a 501-billion-parameter sparse Mixture-of-Experts system that the company says matches China’s best open models on reasoning while consuming a fraction of the inference compute.

What Beam Is

Beam is a sparse Mixture-of-Experts (MoE) model with 501 billion total parameters and 23 billion active per token, built explicitly for coding, reasoning, and agentic workloads. It is a text-only model — no native vision — though Reflection demonstrates it working with information from other modalities when represented as text. The architecture choice is the whole story: by activating only 23B of 501B parameters for any given token, Beam delivers frontier-adjacent capability at a per-token cost profile closer to a small model than a giant one.

The training recipe behind it is unusually well documented for a launch announcement. Reflection pretrained Beam on 23.8 trillion tokens of curated web and proprietary licensed data, which the company says matches or outperforms similar-sized open base models. But the differentiator is reinforcement learning. Reflection ran what it believes is one of the largest RL runs ever conducted by an open lab: more than 100 million rollouts generated on 10,500 Nvidia GB300 GPUs over four weeks, with a maximum context length of 256K tokens during training. Training and grading used approximately 1.3 billion sandboxes, sourced from a pool of nearly one million high-quality coding, agentic, and STEM environments.

The Numbers That Matter

Reflection published a broad benchmark table comparing Beam against the current open-weight field — Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8-Max, and DeepSeek V4.1 Flash. A few results stand out:

  • SWE-bench Verified: 80.9 — ahead of Inkling (77.6) and Nemotron 3 Ultra (70.7)
  • SWE-bench Pro v2-Hard: 77.2 — well ahead of Inkling’s 56.9, though behind Kimi K3 (88.2) and GLM 5.3 (84.3)
  • Terminal Bench v2.1: 80.1 — essentially tied with GLM 5.2 (81.0), within reach of Kimi K3 (88.2) and DeepSeek V4.1 Flash (90.6)
  • AIME 2026: 97.8 — competitive with Inkling (97.1), just under GLM 5.2 (99.2)
  • GPQA Diamond: 90.5 — on the same tier as GLM 5.2 (91.2) and Qwen 3.8-Max (92.6)
  • MCP Atlas: 78.7 — slightly ahead of GLM 5.2 (77.8) on tool calling

The honest picture: Beam is not the raw-capability leader. Kimi K3 and GLM 5.3 beat it on most pure-strength benchmarks, and Reflection says as much — the model “approaches Qwen 3.8-Max” rather than surpassing it. The pitch is efficiency. On advanced reasoning benchmarks, Beam scores comparable to GLM-5.2 while using 3–4× less inference compute, and the gap widens against 2T+ parameter models like Qwen 3.8-Max, which activate vastly more parameters per token. Reflection’s framing: more intelligence per token, translating into a cheaper workhorse for enterprise coding and agentic workloads.

Emergent Behavior Worth Noting

Two technical details from the RL campaign deserve attention beyond the headline numbers.

First, capability transfer. During a training phase focused on reasoning, software engineering, and terminal tasks — with no browsing tasks in the RL mixture — Beam’s browsing performance improved anyway. Given web access, the model organically learned to search for and query other large language models, and to use OCR APIs to read documents. Reflection reads this as evidence that Beam learned generalizable agentic skills rather than benchmark-specific tricks, and demos back the claim: Beam built a live NYC subway dashboard from public MTA data, created a p5.js 3D game, and even produced a fine-tuning notebook that lifted the smallest Gemma-4 model’s accuracy on a held-out Text2SQL test set by 66.5%.

Second, efficient reasoning as a trained behavior. Reflection used a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in RL, performance climbed as completion lengths fell — the model literally learned to solve tasks with less reasoning. Later, longer chains returned, but only when they bought real capability gains. Users get a reasoning effort parameter to control this tradeoff directly, matching depth to task and compute budget.

The Async RL Breakthrough

Under the hood, the most technically interesting claim concerns policy staleness. Fully asynchronous RL lets agents generate rollouts while the trainer learns and publishes new model versions — but tokens from long-running rollouts come from multiple checkpoints, and older tokens drift stale relative to the current policy. Reflection says it developed algorithms that keep learning stable even when training on interactions generated more than a day old — up to 107 weight versions behind the current policy — while systematically closing training-inference numerical mismatch. The infrastructure sustained an average of 110,000 concurrent rollouts. For anyone building agentic RL systems, that section of the technical report may be worth more than the benchmark table.

What Ships, and When

A critical caveat for the impatient: the weights are not downloadable today. Beam is undergoing final red-teaming and evaluations, with early access open via sign-up. Reflection will release the weights, technical report, model card, and developer artifacts “later this month” under an Apache 2.0 license, along with the full stack for running, evaluating, and fine-tuning the model, distributed through hyperscalers and neocloud partners with integrations across open-source libraries and harnesses. Apache 2.0 is the most permissive mainstream license choice — no usage restrictions, no revenue caps — which positions Beam directly against DeepSeek and Qwen releases rather than the “open-ish” Llama-style licenses.

Why This Matters

The stakes are easiest to see in the funding behind it. Reflection raised $2 billion led by Nvidia in October 2025 at an $8 billion valuation, and by March 2026 a further round had pushed it to roughly $25 billion — before shipping a single public model. Beam is the first installment payment on that wager. Nvidia’s interest is structural: every enterprise that downloads open weights and builds its own “AI factory” on self-owned GPUs is a customer Jensen Huang wants more of, and a credible Western open model removes the provenance excuse for security-conscious buyers — banks, defense, government — who won’t run Chinese weights but resent depending on three closed American APIs.

The competitive read is more nuanced than the press release. Beam wins on efficiency and Western provenance; Kimi K3 and GLM 5.3 still win on raw strength; the enterprise decision will come down to cost per resolved task. If Beam’s 3–4× compute advantage holds up in independent deployments, the pricing pressure lands squarely on both closed frontier APIs and Chinese open models alike. And Reflection is explicit that Beam is the first in a series, already training what comes next.

The open-weight frontier just became a two-hemisphere race. The weights land this month; the benchmarks that matter will be the ones users run themselves.

Sources for this report are listed below.