OpenAI’s Jalapeño, the company’s custom AI inference accelerator, took center stage at the Hot Chips conference, where engineers revealed fresh architectural details along with target and real-world performance figures. First shown in June, the chip notably reached tape-out in just nine months. OpenAI says its NUMA-style spatial design allows Jalapeño to outperform Nvidia’s GB200 and GB300 in low-latency inference and in performance-per-watt, while a 2,048-processor configuration scales to 27 exaFLOPS and 32 PB/s of aggregate memory bandwidth.
A Large Chip Built for Massive Scale
Co-developed with Broadcom, Jalapeño is a substantial inference accelerator equipped with 216 GB of HBM4 memory and up to 15.4 TB/s of bandwidth. The processor delivers up to 3.4 MXFP8 PFLOPS and up to 13.4 MXFP4 PFLOPS at a 700W power rating. That makes it well suited to inference workloads, though the MXFP4 format may fall short for training. The silicon is already running in OpenAI’s labs at 1.70 GHz, and engineers said at Hot Chips they plan to raise clocks to 1.80 GHz, likely to push peak performance higher.
Scalability is a central pillar of the design. Jalapeño can link 128 accelerators within a local rack over Ethernet at 600 GB/s, and up to 2,048 ASICs in a 16-rack pod at 200 GB/s per processor. A full system provides 27 EFLOPS of MXFP4 performance, 432 TB of HBM4 memory, and 32 PB/s of aggregate memory bandwidth. Networking relies on Broadcom Tomahawk 6 Ethernet switches paired with a “half-flattened” two-level Clos topology that offers higher bandwidth for tensor-parallel traffic, lower bandwidth for expert-parallel communication, and prioritizes low latency across both. OpenAI confirmed during a Q&A that the scale-up network uses Ethernet with 200-Gb/s links. The physical hardware around the Broadcom-designed chip will be manufactured by Celestica.
On paper, the specifications look modest next to Nvidia’s Blackwell Ultra accelerators, which reach 10 FP8 PFLOPS and 20/15 sparse/dense NVFP4 PFLOPS. OpenAI, however, contends that raw compute and memory bandwidth are not what set its platform apart.
Smart Data Movement Over Raw Bandwidth
According to OpenAI, a 128-ASIC Jalapeño domain delivers more than 1 PB/s of aggregate HBM4 bandwidth, considerably higher than the 576 TB/s of the GB300 NVL72. A one-trillion-parameter model using FP4 weights needs roughly 0.5 TB, meaning the system could in theory read the entire model more than 2,000 times per second. Real-world inference stays well below that ceiling, which is precisely why hardware designers do not add HBM bandwidth without limit.
To close that gap, Jalapeño uses a
Source
Image: tomshardware.com