OpenAI unveiled detailed benchmark results for its Jalapeño AI inference chip at the Hot Chips conference on August 25, 2026. The chip outperforms Nvidia’s GB200 and GB300 superchips by delivering 1.5 to 1.9 times more AI work per watt alongside latency improvements of 1.7 to 3.6 times lower across tested models [1, 2, 3].
Jalapeño is an ASIC designed specifically for AI inference workloads. OpenAI developed it in about nine months in partnership with Broadcom. The chip minimizes delays during the prefill and communication phases of inference by keeping the model state and KV cache local, cutting down on data movement bottlenecks [2, 3].
Richard Ho, OpenAI’s vice president of hardware, said Jalapeño offers the “best of both worlds” with lower latency and higher throughput. He described the performance jump as “a very, very significant advance over state of the art,” noting the chip “can serve more AI work per unit of power, while also returning responses more quickly” [1, 2].
OpenAI introduced Jalapeño in June 2026 and plans to deploy it in small volumes by the end of this year. Larger deployment volumes are expected in 2027 [1, 2]. While the chip is part of a multigenerational platform combining AI models, chips, and memory, OpenAI does not intend for Jalapeño to fully replace Nvidia hardware [1, 2].
Next steps include the initial small-volume deployment of Jalapeño by the end of 2026, with volume ramp-up scheduled for 2027 [1, 2].