NVIDIA Groq 3 LPX Enters Full Production as AI Agents Drive New Inference Architecture
NVIDIA says Groq 3 LPX is now in full production as a specialized inference accelerator for the Vera Rubin platform, targeting low-latency token generation for agentic AI workloads.
Why it matters: Agentic AI is shifting infrastructure competition toward inference latency and cost per token, giving purpose-built decode acceleration greater strategic value.
NVIDIA has moved its Groq 3 LPX inference accelerator into full production, extending the Vera Rubin platform with hardware designed specifically for rapid token generation in agentic AI workloads.
NVIDIA says an Artificial Analysis benchmark running Gemma 4 31B with a 100,000-token context produced about 3,400 output tokens per second. The result is an early benchmark of the production platform and should be read alongside broader independent testing as the hardware reaches more customers.
Nebius is the first AI cloud provider announced for Groq 3 LPX adoption. NVIDIA also says Groq plans deployments with Dell Technologies, while SpaceXAI plans to use NVIDIA Vera CPUs for its next-generation agentic AI architecture.
The architecture reflects a shift in AI infrastructure economics. Training drove the original GPU boom, but autonomous agents can generate long chains of inference steps, making decode latency, throughput, power consumption and cost per token increasingly important.
Groq 3 LPX is NVIDIA's attempt to control more of that inference layer while keeping it integrated with the broader Vera Rubin rack-scale platform.