← All posts / Industry

Hot Chips 2026 Kicks Off at Stanford: OpenAI's Jalapeño Finale and the Inference Silicon Wars

Hot Chips 2026 opens today at Stanford with NVIDIA's Vera and Rubin, Google's TPU 8, Microsoft's MAIA 200, and OpenAI's Jalapeño inference chip closing the show — the clearest sign yet that the AI battle has moved to custom silicon.

Hot Chips 2026 Kicks Off at Stanford: OpenAI's Jalapeño Finale and the Inference Silicon Wars

The most important chip conference of the year starts today. Hot Chips 2026 — the 38th edition of the symposium — runs August 23–25 at Memorial Auditorium on the Stanford University campus in Palo Alto, with in-person attendance sold out. Over three days, the architects of the world’s fastest processors will take the stage to explain, at the gate level, the machines powering the AI boom. And this year’s program reads like a league table of the AI industry itself: NVIDIA, AMD, Intel, Google, Meta, Microsoft, Cerebras — and, in the closing slot on Tuesday, OpenAI.

Sunday: the memory day

Hot Chips opens, unusually, with a day devoted almost entirely to memory — a telling choice in a year when HBM supply is the binding constraint on the entire AI buildout. Micron’s director of HBM design opens the session on where memory architecture goes next. Samsung follows on how the HBM base die evolves once it moves to a leading-edge logic process, and SK hynix covers advanced packaging for HBM — three vendors who between them control the high-bandwidth memory market, each laying out their answer to the bandwidth wall that frontier models keep running into.

Then comes one of the more technically interesting talks of the week: d-Matrix presenting its 3D-DRAM work, first shown at ISCA in June, covering its next-generation AI accelerator Raptor, how 3D DRAM integrates with the tensor engines, and the memory bank architecture that makes it work. The session is co-presented with Meta. Raja Koduri’s startup Oxmiq Labs closes the day with a talk on HBF in AI compute — the memory-fabric technology his company has been championing as an alternative path past the HBM bottleneck.

Monday: NVIDIA’s Rubin, AMD’s MI400, Intel’s trio

Day 1 is the CPU-and-GPU architecture day, and NVIDIA has the first double slot: separate sessions on the Vera CPU and the Rubin GPU — the two headline chips of the Vera Rubin platform that entered full production this summer. Expect deep architectural detail on Vera, NVIDIA’s first fully custom Arm-compatible server CPU (reported at 88 custom cores), and on Rubin, the TSMC 3nm GPU with HBM4 memory that NVIDIA says delivers around 50 petaflops of NVFP4 inference per package via a third-generation Transformer Engine with hardware-accelerated adaptive compression. When NVIDIA presents at Hot Chips, whitepapers follow — these sessions typically become the canonical technical references for the platform.

AMD gets two slots of its own on the Instinct MI400 GPU architecture and the system architecture around it — the rack-scale answer to NVIDIA’s NVL72. Intel, in the middle of its AI-foundry comeback, presents three: Crescent Island, its data-center GPU; Diamond Rapids, the next Xeon; and Wildcat Lake for mobile and PC. For a company raising $20 billion in stock offerings to fund AI ambitions, this is the program where Intel tries to prove the engineering matches the balance sheet.

Tuesday: accelerators, and an OpenAI finale

Day 2 is where the AI accelerator crowd takes over. Norman Jouppi — the Google Fellow who has led TPU development since the first generation a decade ago — presents Google’s eighth-generation TPU family. Announced at Cloud Next in April, the TPU 8 line split into two chips: the training-optimized 8t, which clusters 9,600 chips into a single superpod over a 3D torus network for 121 FP4 exaflops per pod, and the inference-optimized 8i. A full Hot Chips presentation from Jouppi himself means the deepest public look yet at the silicon behind Gemini.

From there, a packed two-hour stretch: Meta on its MTIA AI chip, Microsoft on MAIA 200 (the inference accelerator it has been deploying in Azure), NVIDIA on the Groq 3 LPU — the SRAM-based, deterministic-latency inference accelerator that came out of NVIDIA’s December 2025 licensing deal with Groq and reportedly delivers 35x the inference throughput per megawatt of GPUs — and Cerebras on its wafer-scale approach.

And then the finale: OpenAI. The final presentation of Hot Chips 2026 belongs to Jalapeño, the inference ASIC OpenAI co-developed with Broadcom and unveiled on June 24. The June announcement was striking for two numbers: Jalapeño went from initial design to manufacturing tape-out in nine months — which OpenAI and Broadcom call the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors — and it targets up to a 50% reduction in inference cost per token. What was missing was the technical detail, with a full report promised “in the months after.” A Hot Chips closing slot is exactly where that detail lands. For anyone tracking whether OpenAI can structurally lower its inferenge economics ahead of a potential IPO, Tuesday’s session is the one to watch.

Why this year feels different

Three undercurrents run through the program. First, inference has formally displaced training as the design center of gravity. Google split its TPU line into training and inference chips. Microsoft’s MAIA is “built for inference.” OpenAI’s first-ever chip is an inference chip. NVIDIA licensed an entire inference architecture from Groq. The training race defined 2023–2025; the token-economics race defines what follows.

Second, every hyperscaler and lab now designs its own silicon. Meta, Microsoft, Google, OpenAI, Amazon — all presenting or shipping custom accelerators — while simultaneously remaining some of NVIDIA’s largest customers. Hot Chips has become the place where that tension is quantified, in terabytes per second and tokens per watt.

Third, AI is now designing the chips. The most quoted statistic of the Jalapeño announcement was that the nine-month tape-out was “accelerated by OpenAI models.” If chip design cycles compress from years to months, the cadence of the entire industry — and the durability of every incumbent’s lead — changes with it.

For three days at Stanford, the abstractions get stripped away. No keynotes, no product videos — just architects and their slides. By Tuesday evening, we will know a great deal more about the physical substrate of the next several years of AI. The fire hose starts this afternoon.


Hot Chips 2026 runs August 23–25 at Memorial Auditorium, Stanford University, Palo Alto, with virtual registration available through hotchips.org.