The Benchmark Grows Teeth: MLPerf Inference v6.1 Debuts Vera Rubin, a 512-GPU AMD Cluster, and the First Agentic Tests
MLCommons' biggest inference round ever: 30 submitters, 486 results, 5.7x per-accelerator gains in a year, NVIDIA Vera Rubin NVL72's first peer-reviewed numbers at up to 3.7x GB300, and new End-to-End RAG and Edge Agentic benchmarks.
The industry-standard yardstick for AI inference just had its busiest round ever. On September 16, 2026, MLCommons published the results of MLPerf Inference v6.1, and the numbers tell a story of an ecosystem compounding fast: a record 30 submitting organizations, 486 datacenter and edge results, the largest system ever benchmarked in the suite’s history, and — for the first time — standardized tests for retrieval-augmented generation pipelines and agentic workloads running at the edge.
If you buy, deploy, or simply follow AI infrastructure, this is the one dataset where vendors cannot simply assert superiority. Everything here is peer-reviewed, reproducible, and architecture-neutral. And this round, the pace of change it documents is startling: the best per-accelerator DeepSeek-R1 server result improved 5.7x in a single year (versus the v5.1 round), and the best vision-language-model result improved 2.99x in just six months.
Vera Rubin’s first peer-reviewed numbers
The single most anticipated entry was NVIDIA’s. The Vera Rubin NVL72 rack — the successor to the Blackwell GB300 NVL72 that currently anchors most frontier inference deployments — made its MLPerf debut in the preview category, submitted both by NVIDIA and by cloud provider Nebius on its VR200 NVL72 systems.
The headline figures, drawn from the Closed division where every submitter must run the same reference model for an apples-to-apples comparison:
- Up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL, NVIDIA’s most demanding vision-language benchmark in the suite, achieved across offline, server, and interactive scenarios using vLLM with NVIDIA’s open-source Dynamo inference framework.
- Up to 2.5x higher throughput on DeepSeek-R1 using the TensorRT-LLM library.
NVIDIA attributes the gains to full-stack codesign: enhanced Tensor Cores and a new Transformer Engine accelerating both prefill and decode, NVFP4 precision shrinking the memory footprint of weights, attention, and KV cache, and disaggregated serving with large-scale expert parallelism across the mixture-of-experts layers that models like DeepSeek-R1 and Qwen3-VL depend on. The NVL72’s sixth-generation NVLink scale-up domain delivers 10x higher packet rates and 3x lower latency than commodity Ethernet, which is what makes rack-scale expert parallelism viable at all.
The “preview” designation matters: it means the platform must be commercially available by the next submission round. These are not lab curiosities — they are numbers for silicon that customers will be able to buy within months. NVIDIA also disclosed that in separate preview testing on the SemiAnalysis AgentX agentic benchmark, Vera Rubin NVL72 delivered 30x the performance of GB300 NVL72 — a figure that, while outside MLPerf’s review, hints at how dramatically agentic workloads reward the new architecture’s latency profile.
The largest system in MLPerf history — and it’s AMD’s
Not to be outdone, AMD and first-time submitter Crusoe entered the largest system ever submitted to MLPerf Inference: 512 Instinct MI355X GPUs across 64 nodes on a standard RoCE Ethernet fabric, running gpt-oss-120b and DeepSeek-R1.
The cluster delivered 5.75 million tokens per second offline and 5.39 million in the server scenario on gpt-oss-120b, plus 2.90 million offline and 2.41 million server on DeepSeek-R1 — which AMD calls the highest aggregate token throughput in MLPerf history, with near-linear scaling from 1 to 64 nodes. The gpt-oss run served the model as 512 independent single-GPU replicas in native MXFP4; DeepSeek-R1 used SGLang with one eight-GPU replica per node.
AMD also debuted the Instinct MI350P, a dual-slot PCIe 5.0 card that packages CDNA 4 silicon for standard enterprise chassis: 128 compute units, 144 GB of HBM3E, 4 TB/s of bandwidth, and up to 4.6 PFLOPS of MXFP4/MXFP6 compute — roughly 80% of the flagship MI355X platform’s benchmark performance for deployments with tight power and cooling budgets.
Two new tests, built for how AI actually ships now
The more consequential news may be the benchmark suite itself evolving. v6.1 introduces two tests that mirror real deployment patterns far better than raw single-model throughput:
End-to-End RAG is the first multi-component MLPerf inference benchmark, scoring an entire question-answering pipeline: an embedding model vectorizes the query, a retriever pulls candidate passages from a vector database, a re-ranker refines them, and LLMs reason over the evidence. The reference implementation chains gpt-oss-120B for query decomposition and answer generation, gpt-oss-20B for document grading, e5-base-v2 for embeddings, and ColBERTv2 for re-ranking, over a corpus of 107,484 passages and 824 multi-hop questions from Google’s FRAMES dataset — with a Llama 3.1-8B judge enforcing a 97% accuracy gate.
Edge Agentic Inference targets the coding-assistant pattern now running on workstations and desktop AI boxes. It replays 20 agentic coding trajectories drawn from SWE-bench Verified — 1,007 turns in total — through Qwen3.6-27B as a Q4_K_M GGUF under llama.cpp, single-stream, the way a developer actually runs an agent on a laptop. The metric is mean latency per turn, with a BFCL v4 accuracy gate. First-time submitter Atlas Inference used it to compare an NVIDIA DGX Spark (20.1 tokens/second, all 1,007 turns in under 64 minutes) against an AMD Strix Halo desktop (19.63 tokens/second) — precisely the kind of controlled, same-engine comparison the benchmark was designed to surface.
Heterogeneous, transpacific, and software-defined
Two systems in this round break MLPerf’s traditional mold. Cisco submitted the benchmark’s first cross-vendor heterogeneous system, pooling eight NVIDIA H200 and eight AMD Instinct MI350X GPUs into a single inference pool over a Silicon One G200 fabric — the mixed-accelerator architecture Cisco now sells through its Secure AI Factory. And MangoBoost, with Dell, TensorWave, and Microsoft Azure, ran the first geographically distributed submission: 16 MI300X GPUs in Korea bridged to 16 MI355X GPUs in the United States, serving one GPT-OSS-120B endpoint across the Pacific at a reported 97% scaling efficiency.
Meanwhile, the round is a quiet showcase for software. Intel reports gpt-oss-120B improving 36% on identical Arc Pro B70 hardware, and a 2.4x software-only gain for Llama 3.1-8B on Xeon 6980P. AMD’s ROCm 7 lifted GPT-OSS-120B throughput 28-38% on unchanged MI355X silicon — a 72-GPU cluster now out-throughputs a 94-GPU configuration from the previous round. NVIDIA’s GB300 NVL72 improved up to 1.6x on Qwen3-VL through KV-cache precision cuts, kernel fusion, and disaggregated serving alone.
The end of MLPerf Inference as we know it
Perhaps the most forward-looking detail: 16 of the 30 submitters already used MLPerf’s new API-centric harness, which drives the system under test as a true client/server endpoint over standard APIs — the way inference is actually consumed. That harness is the foundation of MLPerf Endpoints, which opens rolling submissions in October 2026 and will replace MLPerf Inference as the datacenter benchmark in 2027, with normalized results and expanded agentic workloads planned.
For buyers, the v6.1 results are a rare gift: trusted, comparable data across five new accelerators, rack-scale systems from four generations of silicon, and — for the first time — meaningful measurements of the RAG pipelines and agentic loops that constitute real production AI. The benchmark grew teeth this round. The next one intends to swallow its predecessor whole.
Sources
- [1] https://mlcommons.org/benchmarks/inference-datacenter/
- [2] https://www.globenewswire.com/news-release/2026/09/16/3363348/0/en/mlcommons-sets-participation-record-with-new-mlperf-inference-v6-1-benchmark-results.html
- [3] https://www.unite.ai/nvidia-vera-rubin-nvl72-posts-first-mlperf-inference-preview-results/
- [4] https://www.storagereview.com/news/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers
- [5] https://www.nvidia.com/en-us/data-center/resources/mlperf-benchmarks/
- [6] https://www.amd.com/en/blogs/2026/amd-delivers-its-broadest-mlperf-inference-6-1-submission.html