OpenAI publishes first Jalapeño chip benchmarks, claims up to 3.6x lower latency than Nvidia
OpenAI has published the first performance results for Jalapeño, the custom inference chip it developed with Broadcom, and the company says the chip beats comparable Nvidia GB200 and GB300 rack systems on both throughput per watt and response latency.
What's new
In a post on its site, OpenAI wrote: "Across all three, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems." The three workloads tested were GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — a mix of OpenAI's own open model and two rival open-weight models, chosen as public, reproducible benchmarks rather than OpenAI's own production models.
The published numbers are specific: on GPT-OSS 120B, OpenAI reported roughly 85,448 mixed tokens per kilowatt versus 44,960 for the comparison system (about 1.9x), with end-to-end latency of 1.03 seconds versus 1.80 seconds (about 1.7x faster). On DeepSeek R1 670B, the gap widened to roughly 3.6x on latency (1.65 seconds versus 5.99 seconds). On the largest model tested, Kimi K2.5 1T, OpenAI said Jalapeño delivered about 1.5x higher performance per watt and 3.4x lower latency than the comparison system.
OpenAI said the chip was designed to minimize data movement and communication delays, with model state — including the KV cache used while generating a response — explicitly placed and kept local to reduce the overhead that typically comes from moving data between chips during inference.
Context
OpenAI and Broadcom first unveiled Jalapeño in June 2026 as a custom AI chip built specifically for LLM inference, aiming to cut the cost of serving models at scale compared to general-purpose GPUs. This week's post is the first time OpenAI has put concrete performance numbers behind that announcement, roughly two months after the original unveiling. OpenAI said deployment begins in small volumes by late 2026, ramping through 2027 — meaning the chip is not yet in production use at scale.
The release lands alongside a broader push by OpenAI to frame its infrastructure as a single integrated system spanning chips, data centers, and models, and follows a string of hardware moves from Nvidia itself this week, including new Vera Rubin NVL72 throughput claims and expanded NVLink Fusion support for custom accelerators.
Why it matters
Inference cost, not training cost, increasingly sets the economics of running large language models at consumer and enterprise scale, since every user query consumes compute continuously rather than once. If OpenAI's claimed efficiency gains hold up under independent, production-scale testing, custom silicon could meaningfully lower the cost of serving frontier models and reduce OpenAI's dependence on Nvidia GPUs for its highest-volume workloads. That said, these are OpenAI's own benchmarks on OpenAI's own chip, tested against systems the company selected for comparison — not third-party or peer-reviewed results, and not yet a production deployment. The real test will be whether the efficiency gains persist once Jalapeño is running at the volumes OpenAI actually needs.
Corroborating sources
- Openai
https://openai.com/index/jalapeno-first-results/
“Across all three, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems.”
- Techcrunch
https://techcrunch.com/2026/08/25/openais-jalapeno-chip-is-built-for-fast-inference-at-scale-benchmarks-show/