OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the “best of both worlds” with lower latency and higher throughput, as AI systems typically “have to make a trade-off between the two.”
First introduced in June, Jalapeño is an Application-Specific Integrated Circuit (ASIC) made in partnership with Broadcom. It’s designed for AI inference — the process of running a trained AI model to complete a task or deploy an agent.
To measure Jalapeño’s performance, OpenAI used InferenceX, a benchmarking platform that shows how well AI systems handle inference. The test compared Jalapeño’s performance against the best results recorded at the time, which were with Nvidia’s GB200 or GB300 superchips. OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T than the comparison systems, while offering 1.7 to 3.6 times lower end-to-end latency across the three models. That means the chip can provide users with “faster responses, more responsive agents, and more reliable access as the demand grows,” according to Ho.
OpenAI plans to deploy Jalapeño in “small volumes” by the end of this year, but will begin to “ramp the volume up” into 2027, Ho added. The company doesn’t say how many chips it plans to deploy next year, however.
Even with these performance improvements, Ho said OpenAI doesn’t expect to replace its entire chip lineup with Jalapeño, saying its overall compute strategy includes “very good partners,” like Nvidia. OpenAI will continue developing the second and third generations of the new chip.

