TechMeme
Nvidia says its Groq 3 LPX racks delivered 3,400 tokens per second in an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token input sequence (The Register)
Nvidia’s Groq 3 LPX inference accelerator has entered full production, with Nebius signing as its first customer and SpaceX planning to deploy Vera CPUs. In an Artificial Analysis benchmark, the Groq 3 LPX processed Gemma 4 31B on a 100,000-token input sequence at 3,400 tokens per second. The performance demonstrates the capability of Nvidia’s $20 billion investment in Groq’s LPU technology.