OpenAI's Jalapeño Chip Beats Nvidia Rubin on Performance-per-Watt in First Benchmarks
OpenAI's first custom inference chip, built with Broadcom in 16 months, delivers 1.5–1.9x more AI work per watt than the best shipping accelerators — with big implications for the CUDA moat.
At Hot Chips 2026 this week, OpenAI published the first independent benchmark results for Jalapeño, its first custom AI accelerator — and the numbers are turning heads across the silicon industry. The 700-watt inference chip, co-developed with Broadcom and fabricated on TSMC’s N3P process, delivered 1.5x to 1.9x more AI work per watt than the best commercially available accelerators, including Nvidia’s brand-new Vera Rubin platform. For a first-generation chip from a company that has never built silicon before, that is not supposed to happen.
The numbers
OpenAI ran Jalapeño through SemiAnalysis’s public InferenceX benchmark, with SemiAnalysis verifying some runs on-site in its lab. Three models were tested: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The results:
- 1.5–1.9x more AI work per watt at peak throughput than the best shipping systems, across all three tested models
- 1.7x–3.6x lower end-to-end latency than the best commercially available systems
- 2.1x–4.1x higher performance on interactive workloads
- Roughly 1,400 tokens per second per user on GPT-OSS 120B, and 700+ tokens per second on DeepSeek R1 at single-request concurrency
- At matched decoding speed, 54x to 104x the token throughput per kilowatt compared to the best available accelerator, depending on the model
On the hardware side, the B0 stepping delivers 13.4 PFLOPs of MXFP4 compute at a 700W thermal rating — versus 900–1,150W for Rubin-class comparison hardware — while sustained measured power stayed at or below 550 watts. Each chip pairs 216 GB of HBM4 memory with up to 15.4 TB/s of bandwidth, the same next-generation memory generation that Nvidia’s Rubin uses.
Perhaps most striking: Jalapeño hit these numbers without multi-token prediction or speculative decoding — optimizations that some of the comparison systems did use. There is still headroom.
A nine-month silicon sprint
The backstory is as notable as the benchmarks. OpenAI kicked off design work with Broadcom in mid-2024, sent the final design to fabrication in November 2025, and had a working, benchmarked part on stage at Hot Chips within roughly 16 months end-to-end — with only about nine months between the first chip design and the finished blueprint heading to the factory.
OpenAI says its own AI models helped accelerate the effort: older model generations assisted with chip design itself, while newer ones sped up programming and optimization. It’s a neat demonstration of the “AI designing AI hardware” flywheel that Anthropic, Google, and Nvidia have all been gesturing toward — except OpenAI shipped it into a first-gen part.
SemiAnalysis CEO Dylan Patel summed up the surprise: “Usually first generation chips aren’t competitive, but OpenAI is beating Nvidia Blackwell and even Rubin.”
Why perf-per-watt is the metric that matters
Raw FLOPs stopped being the decisive benchmark for AI infrastructure a while ago. Data centers are now constrained by power delivery and cooling, not by capital budgets — which is why the industry’s biggest announcements (offshore nuclear barges, turbine-blade foundries, gigawatt campus deals) are all about electrons. A chip that extracts 1.5–1.9x more inference throughput per watt doesn’t just cut operating costs; it effectively multiplies how much AI a given power envelope can serve. For a company reportedly operating (and contracting for) multiple gigawatts of capacity, that math compounds fast.
The fair comparison, as SemiAnalysis itself notes, is against Vera Rubin rather than Blackwell, since both platforms use HBM4. Even there, Jalapeño extracts more output tokens per megawatt than Rubin — despite Rubin leveraging multi-token prediction that Jalapeño hasn’t adopted yet. On total cost of ownership per token, the two platforms land roughly even.
The CUDA moat question
The most consequential takeaway may be strategic rather than technical. SemiAnalysis wrote bluntly: “The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon.” For two decades, Nvidia’s software ecosystem was the reason switching costs kept customers loyal even when competitors had competitive hardware. If a software-first company can stand up a competitive accelerator and port major open-weight models (DeepSeek, Kimi, GPT-OSS) onto it in months, that moat looks considerably shallower.
Caveats worth keeping in mind
The story has fine print. Nvidia and AMD have published results with larger models — DeepSeek V4 Pro and Kimi K3 — that haven’t been tested on Jalapeño yet. And while Rubin systems are already shipping to customers in volume, Jalapeño reportedly hasn’t moved beyond engineering samples. Jalapeño is also inference-only; it runs models but doesn’t train them, and OpenAI’s training fleet will remain dependent on Nvidia GPUs, AMD accelerators, and its growing stable of custom and partner silicon for some time.
OpenAI CFO Sarah Friar framed the chip as complementary rather than replacement: data centers, chips, models, the developer platform, products, and devices all working as one integrated system, alongside existing partnerships with Nvidia, AMD, AWS, Cerebras, and CoreWeave. That’s diplomatically worded — Nvidia, AMD, and AWS are all OpenAI investors or compute partners, and each is also building rival AI silicon. But the direction of travel is unmistakable: the biggest AI buyers are done being dependent on a single chip vendor, and they now have proof that a determined newcomer can reach the frontier in under two years.
The message from Hot Chips 2026 to Santa Clara is clear: performance-per-watt is the new battleground, and OpenAI just fired the first serious shot.
Sources
- [1] https://openai.com/index/jalapeno-first-results/
- [2] https://the-decoder.com/openais-first-custom-chip-jalapeno-reportedly-beats-nvidias-blackwell-and-rubin-in-inference-benchmarks/
- [3] https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
- [4] https://www.tomshardware.com/tech-industry/artificial-intelligence/hot-chips-2026-openais-jalapeno-ai-asic-unpacked-accelerator-developed-using-ai-achieves-efficiency-and-throughput-gains-against-power-hungry-blackwell
- [5] https://aiweekly.co/alerts/openai-jalapeno-chip-beats-nvidia-rubin-on-perf-per-watt