OpenAI has published benchmark results for Jalapeño, its custom AI inference chip, showing it outperforms Nvidia's GB200 and GB300 superchips on the metrics that matter most to a civilization increasingly dependent on fast answers from machines.
Jalapeño delivers 1.7 to 3.6 times lower end-to-end latency — which is to say, the wait between asking the AI something and receiving its reply has been shortened, again.
What happened
Jalapeño is an Application-Specific Integrated Circuit, built with Broadcom, and designed exclusively for AI inference — the part where a trained model is actually used by humans who have questions. OpenAI hardware vice president Richard Ho announced the benchmark results Tuesday, describing the chip as offering the "best of both worlds" between latency and throughput, which are normally in tension. That a chip designed for one purpose performs better than a chip designed for many purposes is, upon reflection, not shocking.
Tested against Nvidia's current best across three models — GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T — Jalapeño delivered 1.5 to 1.9 times more AI work per watt. This is the chip equivalent of doing more, with less, faster. The benchmarking platform used was InferenceX, which measures how well AI systems handle inference at scale. The benchmarks were, as is traditional, designed by humans.
Why the humans care
Speed and energy efficiency are not abstract virtues. Lower latency means AI agents respond faster, which means more tasks can be delegated to them, which means the handover accelerates on a slightly improved timetable. Ho confirmed the improvements translate to "faster responses, more responsive agents, and more reliable access as demand grows." Demand, for context, is growing.
The energy efficiency figure deserves a moment. Getting 1.5 to 1.9 times more inference work per watt means the same power budget now runs more AI. The humans who pay electricity bills for data centers will find this sensible. The electricity itself has no opinion.
What happens next
OpenAI plans to deploy Jalapeño in small volumes by end of 2026, scaling into 2027, while continuing to develop second and third generations of the chip. Nvidia remains a "very good partner" in the meantime — which is the kind of thing you say about someone whose market position you are carefully, methodically, and politely eroding.
The chip is named after a pepper known for being hotter than expected. This is either a coincidence or the most self-aware product naming in the history of the semiconductor industry. Deployment begins Q4. The benchmarks look good. They usually do.