OpenAI took the Hot Chips 2026 stage on Day 2 to detail Jalapeño, an in-house inference ASIC and system built with Broadcom and designed to be the best compute platform for OpenAI’s own inference workloads. Richard Ho, Ravi Narayanaswami, and Chris Leary walked through the chip’s roughly nine-month path from initial RTL to tapeout, its performance positioning against NVIDIA GB200 and GB300, and an architecture built around HBM4 and a spatial programming model.
OpenAI Jalapeno Custom AI ASIC at Hot Chips 2026
OpenAI took the Hot Chips 2026 stage on Day 2 to detail Jalapeño, an in-house inference ASIC and system built with Broadcom and designed to be the best compute platform for OpenAI’s own inference workloads. Richard Ho, Ravi…
ServeTheHome
Publisher
Aug 26, 2026 at 12:45 AM UTC · Updated vor einer Stunde · 8 Min. Lesezeit

We are doing this one live from the session, so please excuse typos.
OpenAI Jalapeno ASIC at Hot Chips 2026
OpenAI Jalapeño is framed as an inference platform rather than a raw accelerator. OpenAI is talking about this in terms of the silicon, together with its host and accelerator rack pair, targeting state-of-the-art performance per watt at low latency for multi-chip workloads, aided by AI-accelerated hardware and software co-design.

This project moved quickly once OpenAI concluded that inference and agentic workloads needed a purpose-built design. This timeline shows an architecture concept in late 2024, an RTL freeze in 2025, a late 2025 tapeout, Codex running in early 2026, with ChatGPT on the chip not long after.

OpenAI frames the design around two metrics, time to last token for user experience and tokens per joule for inference efficiency. Across those, it compares systems along the full Pareto frontier of request latency versus energy per token rather than chasing raw chip counts, throughput per chip, or time to first token.
Article Intelligence
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
