← All posts / Industry

4.8x on Day One: CoreWeave Puts Nvidia's Vera Rubin NVL72 Into Production With Cognition as First Customer

CoreWeave has delivered Nvidia's Vera Rubin NVL72 at production scale, with Cognition running Devin's training, RL, and inference on the new racks — measuring up to 4.8x total token throughput over GB200 NVL72.

4.8x on Day One: CoreWeave Puts Nvidia's Vera Rubin NVL72 Into Production With Cognition as First Customer

The race to put Nvidia’s next-generation Vera Rubin rack into paying customers’ hands is over, and the winner is not a hyperscaler. At its Fully Connected conference in San Francisco on September 30, CoreWeave announced that NVIDIA Vera Rubin NVL72 is now available on CoreWeave Cloud — with Cognition, the applied AI lab behind the Devin AI software engineer, as the first customer anywhere running production workloads on the system.

What happened

CoreWeave received its first Vera Rubin NVL72 production racks in early September. Within days, and not months, Cognition’s own engineers had stood up a cluster and executed the first customer-run Vera Rubin inference benchmark, measured against a GB200 NVL72 baseline. The numbers they published are striking: up to 4.8x higher total token throughput for SWE-2 inference workloads, and a 3.8x boost in output token throughput for reinforcement learning workloads.

The announcement came alongside a wider set of releases at Fully Connected, which gathered more than 4,500 customers, partners, and developers. CoreWeave also said it will offer NVIDIA Vera, the first CPU purpose-built for AI agents, and launched CoreWeave Forge, a connected environment for training, evaluating, and improving models and agents that unifies Weights & Biases, OpenPipe’s post-training stack, and the open-source marimo notebook project.

Why agentic workloads make rack-scale matter

“Agentic coding is an unforgiving workload that requires long contexts, high concurrency and rapid reasoning,” said Silas Alberti, SVP research and founding team at Cognition. “By deploying the NVIDIA Vera Rubin NVL72 on CoreWeave, our engineers are seeing up to a 4.8 times increase in total token throughput for SWE-2 inference workloads. For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done.”

That compounding effect is the whole story. An agent that solves a software engineering task executes dozens of sequential model calls — plan, read code, edit, run tests, diagnose, repeat. If each step waits on the previous one, raw generation speed improvements multiply across the whole trajectory. For Cognition, the throughput gains translate into more concurrent Devin sessions per GPU, drastically accelerated research loops, and lower cost per session — with, the company says, no loss in generation speed.

The workload Cognition used for benchmarking is itself notable: engineers sampled a subset of tasks from FrontierCode and deployed AI agents to solve them — a real-world software engineering evaluation, not a synthetic latency test. It is the first customer-executed Vera Rubin inference benchmark published anywhere.

Vera Rubin NVL72, briefly

Vera Rubin NVL72 is Nvidia’s rack-scale successor to the GB200/GB300 NVL72 Blackwell generation, pairing the new Vera CPU with Rubin GPUs across a 72-GPU NVL rack connected by NVLink. CoreWeave’s deployment pairs it with Spectrum-X 102.4T Ethernet networking for the scale-out fabric.

CoreWeave has also published the industry’s first measured silicon performance numbers on the platform, claiming 10x token throughput per megawatt over GB200 NVL72 on the DeepSeek R1 reasoning model at matched interactivity. In a market where power, not chips, is increasingly the binding constraint on AI buildouts, throughput-per-megawatt may be the metric that decides which cloud wins the agent era.

The Vera CPU side is aimed at a different bottleneck: sandboxes. Agentic AI pressures infrastructure from two directions — serving agents needs low-latency compute at scale, while improving them through RL requires thousands of isolated environments running simultaneously. A single Vera rack packs 128 CPUs and 11,264 cores, enough for more than 11,000 concurrent one-core environments, hardware-isolated via CoreWeave Sandboxes with Spectrum-X Ethernet switches and BlueField-4 DPUs handling secure agent communication. CoreWeave measured more than 3x faster agent sandbox startup times on Vera CPUs, plus a 1.7x performance gain on Terminal-Bench across passing tasks.

The moat is the operating model

The quieter strategic point in the announcement is continuity. Customers run Vera Rubin under the same operating model and tooling as their existing GB200 NVL72 and GB300 NVL72 fleets — CoreWeave Kubernetes Service, SUNK, Mission Control, Sandboxes, and serverless inference — with CoreWeave’s performance engineering teams layered on top.

“Bringing up NVIDIA Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations,” said Chen Goldberg, EVP of product & engineering at CoreWeave. “Our job is to make compute, networking and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves.”

The CoreWeave–Nvidia relationship dates to 2017 and the Volta generation — and those V100 GPUs are reportedly still running customer workloads nearly a decade later. Nvidia’s Ian Buck, VP of hyperscale and HPC, framed durability as the platform’s core pitch: “infrastructure that keeps earning for years, and the flexibility to put the right GPU on the right workload.” CoreWeave, working with Dell Technologies, was among the first cloud providers to deploy Dell PowerRack systems featuring GB200 and GB300 NVL72, and is one of the first to deploy Vera Rubin.

Analysis: the first-mover economics of new silicon

For CoreWeave, being first to production with each Nvidia generation is the business model. The company has built its identity on record-breaking MLPerf results and its position as the only AI cloud to earn SemiAnalysis’ top Platinum ClusterMAX rating three consecutive times. Early access to new silicon lets it capture the customers whose workloads are most throughput-hungry — and right now, that means agentic AI companies like Cognition, which scaled from bridge capacity to thousands of GPUs for training and inference in under nine months.

For Cognition, the move is a straightforward capacity-and-cost story: Devin’s economics improve nearly five-fold on the throughput axis, and RL research loops — which consume enormous numbers of isolated sandboxes and output tokens — accelerate by almost 4x. In a market where Devin reportedly crossed $1B in annualized revenue, per-session cost is the difference between margin and burn.

The broader signal is about where the agentic AI stack is heading. When the unit of computation shifts from “a chat completion” to “thousands of concurrent multi-step agent trajectories,” the entire platform — GPU fabric, CPU cores for sandboxes, networking, orchestration — becomes the product. Nvidia is designing accordingly (a CPU built for agents, racks built for token throughput per megawatt), and the clouds that co-engineer with it earliest are the ones landing the agent workloads.

Vera Rubin NVL72 is now generally available on CoreWeave Cloud for early-access customers.