← All posts / Industry

OpenAI's First Custom Chip 'Jalapeño' Beats Nvidia on Efficiency in First Public Benchmarks

At Hot Chips 2026, OpenAI showed its Broadcom-built inference ASIC beating Nvidia's GB200/GB300 rack systems with 1.5-1.9x more throughput per kilowatt — from a 700W part.

OpenAI's First Custom Chip 'Jalapeño' Beats Nvidia on Efficiency in First Public Benchmarks

At the Hot Chips 2026 conference, OpenAI arrived with a statement few expected from a first-generation silicon effort: its in-house inference chip, nicknamed Jalapeño, beats Nvidia’s flagship rack systems on the metrics that increasingly define AI data center economics — throughput per watt and token latency.

The benchmarks, presented on August 25 and run using SemiAnalysis’s public InferenceX suite, put Jalapeño at 1.5x to 1.9x more AI work per watt at peak throughput and 1.7x to 3.6x lower end-to-end latency than the best commercially available systems — Nvidia’s GB200 and GB300 — across three open-weight models: GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI’s 1-trillion-parameter Kimi K2.5. For interactive, latency-sensitive workloads, OpenAI claims performance is 2.1x to 4.1x higher, with leads stretching to 8.6x–104.3x more throughput per kilowatt at the GB300’s fastest time-between-tokens settings.

The headline contrast is stark: a 700W ASIC going up against accelerators rated at 1,200W and 1,400W — and winning on efficiency. OpenAI normalized results to each accelerator’s published package TDP, and noted that Jalapeño’s measured sustained power stayed at or below 550W in testing.

What Jalapeño actually is

Jalapeño is OpenAI’s first custom chip, co-developed with Broadcom and fabricated on TSMC’s 3nm-class N3P process. It is an inference-only accelerator — it runs AI models but does not train them — and, notably, it is not tuned exclusively to OpenAI’s own models. It’s a general-purpose LLM inference engine.

Each package pairs its compute die with six HBM4 stacks totaling 216 GiB at 15.4 TB/s of bandwidth. Compare that to the GB300’s 288GB of HBM3E at a 1,400W rating: per watt of rated power, OpenAI’s chip packs roughly 50% more memory. According to OpenAI’s Hot Chips presentation, the core architectural bet is that the bottleneck worth attacking is exposing aggregate HBM bandwidth, not simply adding more of it.

The speed of execution is as remarkable as the silicon. Design work kicked off in mid-2024, the final design went to fabrication in November 2025, and the full cycle took about 16 months — with only nine months between first chip design and tapeout. OpenAI says it used its own AI models during development: older generations helped with chip design, newer ones sped up programming and optimization. A second-generation chip is reportedly approaching tapeout, and concept work on a third generation is already underway.

The numbers, with caveats

OpenAI provided the benchmark numbers, though SemiAnalysis says it verified some runs on-site in OpenAI’s lab. On GPT-OSS, Jalapeño hit roughly 1,400 tokens per second per user; on DeepSeek R1, it topped 700 tok/s on a single concurrent request.

There are important asterisks. Jalapeño achieved these results without multi-token prediction or speculative decoding — optimizations that some comparison systems used, meaning there’s headroom left. An appendix comparison using all-in utility power per accelerator (1.18kW for Jalapeño vs. 2.55kW for GB300) produces narrower gaps, as does pitting Jalapeño against a GB300 running multi-token prediction, where the peak efficiency lead shrinks to roughly 1.5x.

Critically, Vera Rubin was not in the comparison — the Nvidia platform slated to power the first gigawatt of systems OpenAI agreed to deploy in the second half of 2026. SemiAnalysis argues the fairer fight is against Vera Rubin since both use HBM4 memory; even there, Jalapeño squeezes out more output tokens per megawatt, though Nvidia’s part uses multi-token prediction that Jalapeño hasn’t adopted. On total cost of ownership per token, the two come out roughly even. And while Rubin systems are already shipping to customers, Jalapeño reportedly hasn’t moved beyond engineering samples.

Why it matters

SemiAnalysis CEO Dylan Patel summed up the industry shock: “Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin.” The analysis firm went further, writing that “the CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon” — a direct challenge to the assumption that Nvidia’s software ecosystem lock-in would protect it from custom silicon.

The strategic context matters. Nvidia remains deeply entangled with OpenAI’s fortunes: it agreed on August 17 to backstop up to $105 billion in financing for an OpenAI-leased data center campus in Ohio. “Nvidia is a really good partner, and we continue to need a lot of Nvidia,” OpenAI VP of hardware Richard Ho told Bloomberg. CFO Sarah Friar frames Jalapeño as complementing partnerships with Nvidia, AMD, AWS, Cerebras, and CoreWeave rather than replacing them.

But the pressure on Nvidia’s ~75% data-center gross margins is now real. Scaling Jalapeño across the 10GW deployment agreement OpenAI signed with Broadcom last October would also make OpenAI a substantial new claimant on HBM4 supply — a market where Samsung, SK hynix, and Micron have sold capacity through 2027, and where SK hynix’s CEO has warned 2027 will be the worst year of the crunch.

OpenAI plans to begin deploying Jalapeño in its own data centers later this year. For an industry racing to make inference cheap enough for agentic AI workloads that consume tokens by the billions, a 700W chip that outperforms 1,400W flagships per kilowatt is exactly the kind of disruption the incumbents can no longer dismiss as science fiction.