OpenAI's Jalapeño Chip Posts First Benchmarks: Up to 1.9x Nvidia Blackwell Efficiency, 3.6x Lower Latency
OpenAI's first custom inference chip, built with Broadcom, delivered 1.5-1.9x more AI work per watt and up to 3.6x lower latency than Nvidia's GB200/GB300 systems in its first public benchmarks.
Nine months after OpenAI and Broadcom unveiled Jalapeño — the company’s first custom-designed “intelligence processor” for large language model inference — OpenAI has published the first measured performance numbers, and they are turning heads across the semiconductor industry. According to results shared on August 25, 2026, Jalapeño outperformed Nvidia’s flagship Blackwell-generation systems on the two metrics that matter most for inference economics: throughput per watt and end-to-end latency.
The headline numbers
On InferenceX, a public benchmark suite, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems — Nvidia’s GB200 and GB300 racks.
The per-model figures are even more striking. Running the open-weight GPT-OSS 120B model, Jalapeño sustained roughly 1,459 tokens per second per user, compared with approximately 535 tokens per second on Nvidia’s GB200 — a 2.7x advantage in raw decode speed. On the latency side, Jalapeño completed an end-to-end inference task in about 1.65 seconds where the GB300 took nearly 6 seconds, matching the upper end of the 3.6x claim.
OpenAI also reported that, in a matched-latency configuration running DeepSeek R1 decode, Jalapeño delivered up to 104.3x GB300’s throughput per kilowatt. That eye-popping figure comes with important fine print: it reflects a comparison at a specific matched operating point rather than a general-purpose speedup, and analysts were quick to note that the more representative cross-scenario numbers are the 1.5–1.9x perf-per-watt range.
What Jalapeño actually is
Jalapeño is not a GPU. It is a custom ASIC — an “intelligence processor” co-designed with Broadcom and purpose-built for transformer inference at scale. Where a general-purpose GPU must serve training and inference across every customer’s workload, Jalapeño’s silicon is freed to specialize entirely for one customer’s dominant workload: serving large language models.
That specialization shows up in the efficiency curve. SemiAnalysis, whose deep-dive published alongside the results is titled simply “OpenAI Jalapeño: Better Than Nvidia Blackwell,” concluded that Jalapeño beats Blackwell on performance-per-watt “across almost all scenarios without being tuned for any specific point in the curve” — meaning the wins are not cherry-picked benchmark positions but a broad-based efficiency advantage. The analysis also notes that the current B0 stepping of the chip carries roughly a 25% perf-per-watt improvement over the earlier A0 silicon, indicating the design still has headroom as it matures through steppings.
Why this matters
The economics of frontier AI are increasingly inference economics. Training runs are massive but episodic; serving hundreds of millions of users is a permanent, compounding cost. If a hyperscale operator can get even 1.5x more useful tokens per kilowatt — and early third-party estimates suggest total cost of ownership could land anywhere from a third to a fifth of GPU-based serving once the platform is optimized — the savings at OpenAI’s scale run into billions of dollars annually.
It also matters for the industry’s power problem. Datacenter electricity has become the binding constraint on AI expansion, with grid interconnects queued for years in key regions. Chips that do dramatically more work per watt directly relax that constraint, which is why perf-per-watt has replaced raw FLOPS as the metric the industry actually watches.
The caveats
Three deserve attention. First, these are OpenAI’s own measurements on workloads OpenAI selected, on a benchmark (InferenceX) whose configuration favors the workload mix Jalapeño was designed for. Nvidia will certainly contest the framing, and its next-generation platforms are not standing still. Second, silicon leadership is only half the equation — Nvidia’s real moat is CUDA and two decades of software ecosystem; a first-generation custom chip must rebuild that tooling from scratch. Third, and most practically: Jalapeño does not exist at scale yet. Broadcom has indicated small prototype deployments begin around the end of 2026, with the production ramp running through 2027 and full-scale volume beyond that. Nvidia remains the engine of OpenAI’s training fleet and most of its serving capacity for the foreseeable future.
The bigger picture
Jalapeño’s first results confirm a structural shift that has been building for years: the largest AI companies no longer intend to be purely customers of the chip industry. Google built TPUs; Amazon built Trainium; Meta partnered with Broadcom on MTIA; and now OpenAI — Nvidia’s single most famous customer — has silicon that genuinely competes on measured benchmarks. For Nvidia, the risk is not losing all of OpenAI’s business; it is the precedent. Every frontier lab and cloud provider watching these numbers now has proof that a purpose-built inference ASIC from a first-time chip program can beat the incumbent within three years of starting.
For OpenAI, the strategic payoff is leverage. Even if Jalapeño never serves a majority of its traffic, a credible internal alternative reshapes every negotiation over GPU pricing and allocation. In an industry where compute access has been the scarcest resource, that may be worth as much as the watts it saves.
The first real verdict will arrive with volume production in 2027 — when independent labs can buy time on deployed systems and run the benchmarks themselves. Until then, the numbers OpenAI published this week mark the moment the AI chip market officially became a three-way contest between the incumbent, the challengers, and the customers turned competitors.
Sources
- [1] https://openai.com/index/jalapeno-first-results/
- [2] https://openai.com/index/openai-broadcom-jalapeno-inference-chip/
- [3] https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
- [4] https://www.neowin.net/news/openais-upcoming-jalapeo-ai-chip-outperforms-nvidia-gb300-in-inference-tests/
- [5] http://www.aa.com.tr/en/world/openai-says-its-1st-custom-ai-chip-surpasses-nvidia-systems-in-key-tests/4037388