OpenAI's Jalapeño Chip Posts First Benchmarks: Up to 1.9x the Efficiency of Blackwell
At Hot Chips, OpenAI's first custom inference ASIC beat Nvidia Blackwell on SemiAnalysis' InferenceX: 1.5-1.9x more work per watt and 1.7-3.6x lower latency.
At the Hot Chips conference on Tuesday, OpenAI finally put numbers on Jalapeño, its first custom AI inference chip. Tested on SemiAnalysis’ InferenceX benchmark, the ASIC delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than the comparison system — an Nvidia GB200/GB300-class Blackwell deployment. The results were published alongside a detailed engineering blog post and a Hot Chips presentation led by OpenAI hardware chief Richard Ho.
For a company that spent the last two years signing some of the largest compute contracts in history, the message was pointed: OpenAI now builds the most efficient inference silicon it can buy, and it built this silicon itself.
What the benchmarks actually say
The headline numbers cover three very different open-weight models — GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T — spanning mid-size dense models to trillion-parameter mixtures-of-experts. Across all three, Jalapeño simultaneously improved throughput per kilowatt and tokens per user, the two axes that inference providers usually have to trade against each other.
“The bottom line is that the results show a very, very significant performance advance over state of the art,” Richard Ho, OpenAI’s head of hardware, told reporters on a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly.”
That dual win matters more than either number alone. Datacenter economics are increasingly dominated by power, not chip cost: interconnection queues, grid contracts, and cooling capacity are the binding constraints on AI buildouts. A processor that does roughly 50-90% more work per watt at peak throughput doesn’t just cut the electricity bill — it effectively multiplies how much inference a given power envelope can host. And the latency side, measured as time-between-tokens and end-to-end response time, is what makes agentic workloads feel instantaneous instead of sluggish.
The hardware behind the numbers
The DCD-reported Hot Chips details fill out the spec sheet. Each Jalapeño package runs at a 700W TDP with 15.4 TB/s of HBM4 memory bandwidth. A 128-chip deployment reaches 1.7 exaflops of 4-bit compute with 27.5 TB of HBM4 in total. SemiAnalysis notes that the 15.4 TB/s figure implies the HBM4 stack is running at 10 Gbps pin speeds — slightly ahead of the 9.6 Gbps that Nvidia is achieving on comparable stacks, a meaningful edge for memory-bound decoding phases.
The architecture itself is a chiplet design: one large compute die surrounded by HBM memory chiplets and an I/O chiplet, built on TSMC’s 3nm process in partnership with Broadcom. When the chip was first unveiled in June, OpenAI also highlighted roughly 44 GB of on-chip SRAM delivering around 21 PB/s of on-package bandwidth — scratchpad capacity that lets Jalapeño keep hot model state local instead of shuttling it across a network.
That last point is the core of the design philosophy. Inference is not one workload; it is at least three phases — prefill, communication, and decode — each with different bottlenecks. General-purpose GPUs must compromise across all of them. Jalapeño was instead designed, in OpenAI’s words, to “minimize data movement and communication delays”: the KV cache used while generating a response can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each phase. In other words, OpenAI threw a decade of hyperscaler scheduling wisdom at the problem, then baked it into the silicon.
Context: why OpenAI is building chips at all
Jalapeño was first announced in June 2026 as a joint project with Broadcom, with OpenAI’s own models reportedly assisting in the development process. It is part of a broader “full stack behind abundant intelligence” strategy in which products, models, chips, and memory are designed in concert, as a multigenerational platform rather than a one-off part.
The strategic logic is straightforward. Inference is becoming the dominant share of AI compute as chatbots, coding agents, and autonomous workflows scale. Owning an inference-optimized ASIC gives OpenAI leverage on cost per token, independence from GPU allocation cycles, and the ability to co-design future models around a known hardware target — the same flywheel Google has run with TPU for a decade.
The competitive context is equally sharp. Anthropic is developing custom silicon for Claude, Meta is preparing to manufacture its next AI chip in September, and Microsoft and Google continue to expand their own accelerator lines. Nvidia, for its part, is not standing still: the Blackwell generation that Jalapeño just beat on efficiency-per-watt will itself be superseded, and by the time Jalapeño reaches volume deployment the baseline will have moved. Ho was candid about this — and about the fact that OpenAI isn’t abandoning its GPU suppliers, describing the overall compute strategy as still including “very good partners” like Nvidia.
Deployment timeline
Jalapeño will deploy “in very small volumes” at the end of 2026, ramping to significant deployment through 2027. OpenAI did not disclose unit volumes. Second and third generations of the chip are already in development, consistent with the multigenerational platform commitment.
That timeline tempers the benchmark enthusiasm. Impressive results against today’s Blackwell systems must survive contact with Nvidia’s next generation — and with whatever TSMC process economics look like in 2027. But the direction is clear: the era in which frontier AI companies were purely Nvidia’s customers is ending. OpenAI is now a silicon vendor in its own right, and the first customer is itself.
What to watch
Three things will determine whether Jalapeño’s paper advantage becomes structural: real-world fleet utilization versus benchmark conditions (InferenceX results reflect controlled configurations); the pace of Nvidia’s response in the Rubin generation; and whether OpenAI’s full-stack co-design loop — models shaped for the chip, chip shaped for the models — actually compounds across generations. The first data point, delivered today, is a strong one.
Sources
- [1] https://openai.com/index/jalapeno-first-results/
- [2] https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/
- [3] https://www.theverge.com/ai-artificial-intelligence/984290/openai-jalapeno-ai-chip-benchmarks
- [4] https://www.datacenterdynamics.com/en/news/openai-details-jalape%C3%B1o-ai-chip-with-700w-tdp/
- [5] https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia