2.7x Faster MoE Training, Fully Open: Ai2 Ships Olmo-core 3 Into the Trillion-Parameter Era
The Allen Institute's redesigned open training stack keeps experts GPU-resident, hits 858 TFLOP/s per B300, and has been benchmarked past 2.38 trillion parameters.
On October 1, 2026, the Allen Institute for AI (Ai2) released Olmo-core 3, a ground-up redesign of its open training framework for large language models, built around a new mixture-of-experts (MoE) training system. The numbers are striking: roughly 2.7x the training throughput of its predecessor on the same hardware, benchmarks past one trillion total parameters, and a short-capacity test that reached 2.38 trillion. All of it is fully open — code, technical report, and an interactive demo included.
In a frontier-AI economy where training infrastructure is usually the most closely guarded secret a lab owns, Ai2 is deliberately publishing the machinery behind its next model. Olmo-core 3 is described as one of the core systems behind the next generation of Olmo, which Ai2 says will use an MoE architecture and aims to be its most capable Olmo yet — trained on its largest dataset with its longest context window to date.
Why MoE training is hard
Mixture-of-experts models offer a compelling trade: a model can contain far more learned parameters without requiring every input to activate all of them. Only a handful of experts — the specialized components inside the model — fire for each token the model processes. That sparsity is what lets a modern MoE deliver frontier-ish quality at a fraction of the per-token compute of a dense model of equal total size.
The catch is that the full model still has to live somewhere. Every expert must be stored in GPU memory across the cluster and updated during training, and every token must be routed to the right experts, wherever they physically sit. As MoEs grow, that storage and coordination overhead can erode most of the computational advantage that sparsity was supposed to buy. Ai2 puts the problem plainly: as MoEs grow, those costs “can erode much of the computational advantage of using only part of the model for each input.”
One of the release’s most telling benchmarks quantifies exactly this pain point. Ai2 grew the expert pool from 8 experts to 128 while still selecting only four experts per token, keeping active parameters per token roughly fixed at about 3.2 billion. Total parameter capacity jumped from 4.6B to 47B — a tenfold increase in model capacity — while training throughput fell by less than 5%. Capacity became nearly free; the routing fabric absorbed the shock.
From reshuffling weights to resident experts
The heart of the redesign is a change in how the model’s weights live on hardware. Ai2’s earlier MoE implementation in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and re-shard model weights for each micro-batch of training data. Every batch meant an expensive round of collecting scattered weight shards into place, computing with them, and scattering them back.
Olmo-core 3 switches to a system based on distributed data parallelism (DDP) that keeps experts resident on the GPUs and routes the data to them instead. The weights stay put; the tokens travel. Eliminating that repeated weight gathering is the single biggest source of the headline speedup. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, versus 19,400 with the earlier FSDP-based implementation — about 2.7x the throughput, on identical model and hardware.
Ai2 is careful to acknowledge that NVIDIA’s Megatron-Core is an established option for training large MoEs. The point of Olmo-core 3 is not that incumbents are slow — it’s that the framework behind Olmo now ships an integrated, fully open MoE training stack of comparable ambition, one researchers can inspect, fork, and adapt to their own hardware rather than treat as a black box.
Three ways to split a model
Olmo-core 3 combines three techniques for deciding how the model and its training state are divided across a cluster:
- Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool.
- Pipeline parallelism splits the model’s layers — the successive stages that transform an input — across groups of GPUs, reducing how much of the model each GPU must hold in memory.
- A distributed optimizer spreads optimizer state (the additional data used to calculate and apply updates during training) across GPUs instead of storing a full copy on every one.
Together these allow an MoE to scale without requiring any single GPU — or any single copy of the model — to hold everything.
The release also attacks the cost of routing tokens to experts and running their computations:
- Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the extra work of rearranging it. The technical report describes an NVSHMEM-based design that writes each token directly into its expert’s buffer.
- GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back over the PCIe bus.
- Grouped GEMM combines many small expert computations into larger matrix multiplications the GPU can execute far more efficiently — the report notes device-scheduled grouped matrix multiplies make expert parallelism fully synchronization-free.
MXFP8: a 21% speedup and leaner memory
Olmo-core 3 also supports MXFP8, a lower-precision number format that represents some values with fewer bits, cutting both computation and the volume of data moved between GPUs. Ai2 measured its effect in a controlled benchmark on four NVIDIA B300 GPUs with work distributed uniformly across experts. With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than the BF16 baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and data movement between experts, not attention alone.
The report is refreshingly honest that precision tricks carry conversion costs — savings only materialize when they outweigh the cost of converting between number formats, and Olmo-core 3 exposes control over where those trade-offs are made.
Into the trillion-parameter range
At the top end, Ai2 benchmarked Olmo-core 3 across a range of configurations on NVIDIA B300 GPUs, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs, where the highest observed throughput was 858 TFLOP/s per GPU — a measure of useful model computation per second on each GPU. The technical report lists measured operating points from a 12.9-billion-parameter model on 16 GPUs all the way up to that 1.2T model on 512.
Ai2 also experimented with DeepEP v2, an alternative backend for inter-expert communication, reaching a configuration with 2.38 trillion total parameters — though it describes that as a short-capacity test demonstrating reachable scale, not sustained training performance. Importantly, Ai2 states clearly that these headline rates are systems references measured under random routing to gauge machinery performance, not a controlled scaling curve or a time-to-quality claim.
The findings that didn’t work
Perhaps the most valuable part of the release for researchers is the catalog of negative results — things Ai2 tested and, crucially, did not adopt. These rarely make press releases but routinely consume months of engineering time elsewhere:
- Token gerrymandering: a score intended to encourage balanced routing could improve even as the actual workload became less balanced — the metric was lying.
- Lowering experts’ learning rates because they process fewer tokens did not improve results in the model family tested.
- GPU computations took different amounts of time depending on the values being processed, even with identical matrix dimensions — a reminder that real hardware is messier than roofline models.
- Overlapping communication and computation on separate GPU streams did not always help and in some tests slowed end-to-end execution.
- CPU offload of activations was tested but abandoned — the host link could not carry the traffic.
The report also documents topology-agnostic checkpoints that record global FP32 tensors independently of the parallel layout, so a run can resume on a different cluster topology than it started on — a small feature with outsized practical value for anyone who has ever lost a week-long training run to a hardware shuffle.
Why it matters
Frontier training stacks — Megatron-Core, NeMo, JAX-based systems like GSPMD — are powerful but largely opaque to outsiders, and the fully open alternatives have historically topped out well below trillion-parameter scale. Olmo-core 3 changes that reference point: an academic-adjacent lab has now published a working, benchmarked, trillion-parameter-class MoE training stack, warts (random-routing caveats, negative results) included.
For academic researchers and smaller labs, the release directly attacks the barrier Ai2 names: training large models “takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many.” A 2.7x throughput gain on the same GPU budget is, functionally, a 2.7x cut in the cost of every experiment. For the broader ecosystem, it means the next generation of Olmo — and whatever the community builds on top of this stack — will be reproducible in a way that closed frontier models never are.
The technical report, “Supercharging Olmo-core for Efficient and Scalable MoE Training,” is dated October 2026, authored by researchers from the Allen Institute for AI and the University of Washington, with Tianhua Tao as lead author and primary developer. The code is open now; the next Olmo, Ai2 says, will be built on it.