OpenAI Jalapeño, the company’s new AI chip, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the “best of both worlds” with lower latency and higher throughput, noting that AI systems typically “have to make a trade-off between the two.”
First introduced in June, Jalapeño is an Application-Specific Integrated Circuit (ASIC) developed in partnership with Broadcom. It is designed for AI inference, the process of running a trained AI model to complete a task or deploy an agent.
How Jalapeño performed in benchmarks
To measure Jalapeño’s performance, OpenAI used InferenceX, a benchmarking platform that shows how well AI systems handle inference. The test compared Jalapeño against the best results recorded at the time, which came from Nvidia’s GB200 and GB300 superchips.
OpenAI says Jalapeño delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1 and Kimi K2.5 1T than the comparison systems, while offering 1.7 to 3.6 times lower end-to-end latency across the three models. That means the chip can provide users with “faster responses, more responsive agents, and more reliable access as the demand grows,” according to Ho.
Deployment timeline and strategy
OpenAI plans to deploy Jalapeño in “small volumes” by the end of this year, but will begin to “ramp the volume up” into 2027, Ho added. The company has not said how many chips it plans to deploy next year.
Even with these performance improvements, Ho said OpenAI does not expect to replace its entire chip lineup with Jalapeño, adding that its overall compute strategy includes “very good partners,” such as Nvidia. OpenAI will continue developing the second and third generations of the new chip.
Source
Image: theverge.com