Skip to content
Computing

Nvidia Unveils Groq 3 LPX Architecture at Hot Chips 2026

At Hot Chips 2026, Nvidia presented the architecture behind its Groq 3 LPX inference rack and released the first independent benchmark of the hardware. The presentation was delivered by Igor Arsovski, now Nvidia's VP of hardware and the former chief architect at Groq, whose inference chip has...

Nvidia Unveils Groq 3 LPX Architecture at Hot Chips 2026
At Hot Chips 2026, Nvidia presented the architecture behind its Groq 3 LPX inference rack and released the first independent benchmark of the hardware. The presentation was delivered by Igor Arsovski, now Nvidia's VP of

At Hot Chips 2026, Nvidia presented the architecture behind its Groq 3 LPX inference rack and released the first independent benchmark of the hardware. The presentation was delivered by Igor Arsovski, now Nvidia’s VP of hardware and the former chief architect at Groq, whose inference chip has become Nvidia silicon following the two companies’ deal.

The rack is built on the LP30 chip, which Nvidia acquired through its EUR 17 billion Groq acquisition in December 2025. That same transaction pushed the Rubin CPX, which the LP30 replaces, off Nvidia’s roadmap. Arsovski said the LPX rack is already in production.

Independent testing firm Artificial Analysis measured the rack at 3,431 output tokens per second on a 100K-context Gemma 4 31B reasoning workload, roughly four times the 870 tokens per second posted by the next-fastest public endpoint. The comparison ran on a private, pre-release Gemma 4 31B endpoint served through Google Cloud, using the median of 50 sequential client requests at a concurrency of one. The public providers measured against it ran shared production serverless endpoints, so the single-request setup that produces the highest per-user token rate is not directly comparable to multi-tenant conditions.

Nvidia’s on-stage demo showed a higher figure of 10,996 tokens per second on the same 31B model, which Arsovski flagged as “self-reported,” adding that the goal was “third-party verified independent benchmarks that you guys can trust.” Gemma 4 31B is a dense model small enough to fit inside a single LPX rack, and performance at trillion-parameter mixture-of-experts scale, where memory capacity becomes the main constraint, was not addressed.

An SRAM Design Without HBM

Each LP30 carries roughly 500MB of on-die SRAM and no HBM. A full LPX rack of 256 chips holds 128GB of memory delivering 40 PB/s of aggregate bandwidth against 315 PFLOPS of FP8 compute, with 350 ns of chip-to-chip latency. The Vera Rubin-compatible, MGX liquid-cooled rack scales past 1,000 LPUs.

Keeping model weights resident in SRAM rather than streaming them from HBM removes the memory-access latency that dominates single-token decode. The design drops caches, branch prediction, and out-of-order execution in favor of a fully deterministic pipeline that the compiler schedules at clock-cycle granularity. The architecture descends from the Tensor Streaming Processor that Groq, founded by former Google TPU engineer Jonathan Ross, described in a 2020 ISCA paper titled Think Fast, the same title reused at Hot Chips.

Capacity Is the Trade-Off

A Rubin GPU carries 288GB of HBM4, roughly 576 times the memory of a single LP30. As a result, a 31-billion-parameter model at FP8 needs on the order of 62 LPUs to hold its weights, and a large mixture-of-experts model runs into four figures of chips spread across several racks. Capacity is the cost of the SR

Source
Image: tomshardware.com

The US tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your US list is ready.

Shop Amazon Tech Deals Shop Amazon Tech Deals