← All posts / Industry

Raptor Meets the AI Factory: d-Matrix Plugs Its 3D-DRAM Inference Chips Into NVIDIA's Racks

d-Matrix will ship its Raptor XPUs inside NVIDIA MGX racks via NVLink Fusion, pairing memory-centric inference silicon with Vera CPUs and Spectrum-X networking for premium low-latency token services from Q4 2027.

Raptor Meets the AI Factory: d-Matrix Plugs Its 3D-DRAM Inference Chips Into NVIDIA's Racks

Seven years after founding d-Matrix on the bet that inference — not training — would become the dominant cost in AI computing, CEO Sid Sheth has signed the deal that puts his chips inside the buildings he once had to fight to enter. On September 10, 2026, d-Matrix announced a multi-year collaboration with NVIDIA: its next-generation Raptor inference XPUs will be integrated directly into NVIDIA’s MGX rack-scale architecture using NVLink Fusion, with initial availability expected in Q4 2027.

The announcement, made via PR Newswire and previewed on NVIDIA’s newsroom, turns a would-be NVIDIA competitor into a card-carrying member of the NVIDIA AI factory ecosystem — and it is the clearest signal yet that the AI inference market is splitting into a layered, heterogeneous stack rather than a winner-take-all GPU market.

What was announced

At the core of the deal is an NVLink Fusion-enabled rack-level system that AI labs, hyperscalers, and neocloud providers can deploy for what the companies call “premium-level token services.” Rather than building its own liquid-cooled rack and fighting for datacenter floor space on its own, d-Matrix will plug Raptor XPU trays into the same NVL144 MGX rack architecture that NVIDIA already ships around the world.

The integrated system reads like a full NVIDIA family portrait: NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet networking, with d-Matrix Raptor XPUs sitting in the compute trays where Vera-Rubin GPUs would otherwise live. A single rack will hold 144 Raptor XPUs. Connectivity partner Astera Labs — itself a key player in the NVLink Fusion ecosystem — will deliver the custom connectivity silicon that keeps data flowing between the d-Matrix trays and the rest of the rack.

For d-Matrix, the economics of joining rather than competing are straightforward. “This collaboration with NVIDIA is a defining moment on our journey to infinite inference, accessible to all,” Sheth said in the announcement. “Being integrated into NVIDIA’s latest MGX rack-scale infrastructure with NVLink Fusion means our customers can deploy our inference XPUs alongside the broadly available NVIDIA AI factory platform.”

Jensen Huang framed it as platform strategy rather than mercy toward a rival: “NVLink Fusion enables partners to integrate custom silicon with NVIDIA’s deep ecosystem of NVLink, advanced packaging, rack-scale systems and networking technologies. With NVIDIA AI infrastructure deployed across cloud and on-premises data centers worldwide, NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms — expanding accelerator choice for customers building the next generation of AI factories.”

Why memory-centric inference silicon?

d-Matrix was founded in 2019, “when not many people even knew what inference was,” as Sheth put it in a briefing with journalists and analysts this week. The company’s core thesis is memory-centric computing: instead of shuttling weights between discrete DRAM and a GPU’s compute units, d-Matrix fuses SRAM and 3D-stacked DRAM directly with compute. The result, the company claims, is inference workloads that run up to ten times faster than GPU-only deployments at a fraction of the energy.

The timing of this deal is no accident. Over the past twelve months, agentic AI tools — Anthropic’s Claude Code, coding agents, voice assistants, real-time chatbots — have pushed low-latency inference demand “off the charts,” in Sheth’s words. Tokens-per-second per user is becoming a product feature, not a backend metric. Cybersecurity applications, where response latency directly translates to blast radius, are adding yet another wave of demand.

The first product, Corsair, entered full production in June 2026 after more than $500 million in cumulative funding, including investment from Microsoft’s M12 venture arm. Raptor is the follow-on, and it is the chip this NVIDIA deal is built around.

Inside Raptor: a “two-story” chip

Raptor extends the memory-centric architecture with a first-of-its-kind 3D DRAM stacking approach that brings a DRAM memory chip and an SRAM compute chip together into a single “two-story” package. The technical details, published by the IEEE and previewed by d-Matrix co-founder and CTO Sudeep Bhoja at Hot Chips 2026, describe a 4-nanometer TSMC compute die fused atop a custom DRAM die at a 36-micron pitch, delivering 100 TB/sec of memory bandwidth — the spec that makes ultra-low-latency decode possible.

Per the Next Platform’s reporting, a rack of Raptor chips delivers 2.3 TB of 3D-stack DRAM per card, an aggregate memory bandwidth of 7.2 petabytes per second per rack, and the option to run as a companion rack alongside NVIDIA Vera-Rubin GPU racks. At Hot Chips, d-Matrix showed Z.ai’s GLM 5.2 running at roughly 3,000 tokens per second per user and Moonshot AI’s Kimi K3 at 1,000 tokens per second per user, with the vendor claiming scalability across eight racks.

The division of labor in a heterogeneous deployment is the actual product: GPUs excel at the compute-intensive prefill phase, where context is processed, while Raptor accelerates the latency-sensitive decode phase, where the model streams out its response. Operators using heterogeneous disaggregation can split workloads between d-Matrix Raptor XPUs and NVIDIA Vera Rubin, optimizing each phase of inference independently.

The premium token economy

The commercial logic behind the partnership is the emergence of a two-tier token market. Commodity tokens — batch analytics, background processing — can run anywhere cheap. Interactive tokens — coding assistants that autocomplete in real time, voice agents that must respond in conversational rhythm, chatbots where every 100 milliseconds of latency is felt — are worth a premium.

That premium tier is what d-Matrix and NVIDIA are jointly targeting. NVIDIA’s AI factory business is expanding rapidly: CFO Colette Kress said in the company’s Q2 2027 earnings that the revenue opportunity per gigawatt has grown from about $18 billion in the Grace-Hopper era to roughly $40 billion per gigawatt with the forthcoming Vera-Rubin systems. d-Matrix is effectively signing up to take a slice of that growing pie from inside the pie.

Jesse Clayton, principal product marketing manager for NVIDIA’s Data Center GPU business, described the architecture as “vertically integrated, but horizontally open” — partners can use as much or as little of the NVIDIA stack as they want, plugging in their own silicon through NVLink Fusion ports. The platform, in his words, is “completely fungible.”

Analysis: co-opetition becomes the official strategy

The most significant aspect of this deal is what it says about NVIDIA’s strategy — and about the limits of challenging NVIDIA head-on.

For NVIDIA, every NVLink Fusion partner is simultaneously a competitor neutralized and a moat reinforced. d-Matrix’s XPUs will ship inside NVIDIA-designed racks, connected by NVIDIA interconnect, managed within NVIDIA’s AI factory software, and sold into NVIDIA’s customer base. NVIDIA captures the ecosystem revenue regardless of whose silicon wins a given workload. It is the same playbook the company has run with MediaTek in edge computing, and it mirrors the GPU giant’s July 2026 partnership with d-Matrix on a joint inference system that today’s announcement extends into a formal multi-year roadmap.

For d-Matrix, the trade is distribution in exchange for dependence. Building its own racks (via its April acquisition of GigaIO’s datacenter business and the SquadRack reference design built with Arista, Broadcom, and Supermicro) was always an option, as Sheth acknowledged — Raptor “needs a house to put it in.” But NVIDIA’s house is already deployed in every major datacenter on Earth, with a proven liquid-cooled supply chain and an operations playbook that hyperscalers trust. “We take the Raptor trays, plug them into the same NVL144 MGX rack architecture, which is widely deployed across many datacenters, and we get instant access to those datacenters,” Sheth told The Next Platform.

For the broader market, the deal formalizes what operators have been signaling for a year: inference is becoming heterogeneous. The Rack now matters more than the chip, and the winners will be those who can mix compute types per phase of workload. Whether Raptor’s Q4 2027 availability — more than a year away — lands in a market that has moved on to other memory-stacking schemes (Ferrocompute’s ferroelectric approach, Positron’s memory-first accelerators, Kepler Computing’s compute-in-memory) is the open question. But with backing from more than 100 patents, active evaluation at hyperscalers and frontier labs, and now NVIDIA’s own rack architecture as a distribution channel, d-Matrix has bought itself a seat at the table it was built to disrupt.