OpenAI and Broadcom unveil Jalapeño, a custom LLM inference chip with 50% lower cost than GPUs
OpenAI and Broadcom on June 24, 2026 jointly announced Jalapeño, a custom AI accelerator built specifically for large language model inference. Developed in nine months — what the companies call one of the fastest ASIC development cycles ever achieved in high-performance semiconductors — the chip is the first result of a multi-generation compute partnership between the two companies.
What's new
Jalapeño is designed from scratch around OpenAI's understanding of LLM inference workloads. Unlike general-purpose GPU clusters, it targets the specific compute patterns of real-time model serving, including coding-focused models where operating cost is a primary constraint. Key specifics announced at launch:
- Development pace: Initial design to manufacturing tape-out completed in nine months, with Broadcom and Celestica industrializing the platform.
- Performance: Early testing shows substantially better performance-per-watt than current state-of-the-art GPU alternatives, according to OpenAI.
- Cost: Broadcom CEO Hock Tan cited cost savings of roughly 50% compared with typical AI GPUs.
- Scope: Optimized for inference and real-time serving. Pre-training workloads will continue to run on NVIDIA hardware.
- Timeline: Initial deployment targeted for end of 2026, with expansion in subsequent years.
OpenAI President Greg Brockman framed the chip as a natural consequence of the company's vertical integration: "We have a deep understanding of the workload...how can we build something that will accelerate what's possible?"
Context
OpenAI has long run its inference workloads on NVIDIA GPUs, as has virtually every other frontier AI lab. The company has been public about the enormous compute costs involved in serving models at scale, particularly as usage of GPT-5-class models has grown. Rumors of a custom silicon effort surfaced over the past year, but today marks the first official confirmation and public demonstration of the program.
The partnership with Broadcom places OpenAI alongside Google (TPU), Amazon (Trainium/Inferentia), and Microsoft (Maia) in the cohort of hyperscalers building custom AI accelerators. Broadcom has an established track record in custom ASIC design for cloud customers, making it a logical manufacturing partner.
Why it matters
Jalapeño is primarily an economics story. If 50% cost reductions on inference hold at scale, that directly expands the margin on ChatGPT and API products, which are OpenAI's primary revenue sources. Lower inference costs also make it feasible to serve more demanding real-time applications — voice, coding agents, complex agentic loops — at competitive price points.
The chip also signals a strategic shift. Building custom silicon requires multi-year roadmap commitments, deep engineering investment, and a willingness to bet on volume. That OpenAI is doing this now, at the scale implied by "gigawatt scale data centers with Microsoft," suggests the company is planning for a sustained period of high-volume, cost-sensitive serving demand — consistent with an IPO narrative built around profitable AI infrastructure.
For NVIDIA, the announcement confirms that inference — already under competitive pressure from AMD and custom silicon at Google and Amazon — now faces a direct OpenAI-Broadcom challenger. Pre-training remains a NVIDIA stronghold for now, but inference is where the majority of GPU cycles are consumed in a deployed product.
Corroborating sources
- Techcrunch
https://techcrunch.com/2026/06/24/openai-unveils-its-first-custom-chip-built-by-broadcom/
“Because OpenAI operates across the stack, each layer can be optimized around the same goal: making its models faster, more reliable, and more affordable.”
- Bloomberg
https://www.bloomberg.com/news/articles/2026-06-24/openai-and-broadcom-unveil-ai-chip-to-run-models-faster-cheaper