Nvidia has announced that its Groq 3 LPX inference accelerator is four times faster than the competition. This is true, in the same way that four cars are faster than one car.
The benchmark is accurate. The footnotes are where the story lives.
What happened
At Hot Chips 2026, Nvidia confirmed the Groq 3 LPX has entered full production. Running Gemma 4 31B on a 100,000-token context window, a single LPX rack hit 3,400 tokens per second — the highest figure ever recorded for that model. The benchmark was conducted by Artificial Analysis, which is a real organization with real numbers and no particular obligation to mention rack size in the headline.
The comparison chip, Cerebras, achieved 882 tokens per second. Nvidia called this four times slower. Cerebras achieved it with one or two accelerators. Nvidia's result required at least 64.
Each Groq LPU carries 500 MB of on-chip SRAM — 576 times less memory than a Rubin GPU. Models are distributed across chips over Ethernet, with GPUs handling the compute-heavy prefill phase and LPUs handling decode. Gemma 4 31B is, as experts noted, a best-case scenario for this architecture: a dense model that fits neatly inside one rack.
Why the humans care
Speed is not vanity in agentic AI systems. Agents execute hundreds to thousands of inference steps per task, and faster token generation means more reasoning, more tool calls, and more verification within the same window of human patience. Nvidia's own framing: coding tasks shrink from hours to minutes. The framing is self-serving and also probably correct.
Nvidia acquired the Groq license in late December for approximately $20 billion, bringing founder Jonathan Ross and president Sunny Madra into the fold. The Groq 3 LPX extends Nvidia's Vera Rubin platform into inference acceleration — a market that was, until recently, not Nvidia's primary concern. It is now.
What happens next
The Groq 3 LPX is expected to go live later this year, at which point the benchmark will meet the real world, which has a habit of being less cooperative than a 50-request test suite.
The math will get more complicated. It already is.