NVIDIA's Groq 3 LPX inference accelerator enters full production
NVIDIA announced on August 24, 2026 that Groq 3 LPX, its dedicated interactive AI inference accelerator, is now in full production. The chip is built to speed up the "decode" phase of inference — the stage that determines how fast tokens stream back to a user — and Nebius has signed on as the first AI cloud provider to deploy it.
What's new
NVIDIA's own announcement states plainly: "NVIDIA Groq 3 LPX, the interactive AI inference accelerator, is now in full production." The company positions it as "an extension of the NVIDIA Vera Rubin platform," built to enable "ultrafast token generation for highly responsive agentic systems."
On performance, NVIDIA cites third-party benchmarking: the system "delivered a record 3,400 output tokens per second in Artificial Analysis benchmarking running Gemma 4 31B," and NVIDIA claims "4x faster responsiveness for agents and latency-sensitive workloads than the nearest alternative platform." Racks package 256 individual Groq 3 chips together, according to reporting on the announcement.
On adoption, NVIDIA says "Nebius is the first AI cloud to adopt NVIDIA Groq 3 LPX," with plans to fold it into "Nebius Token Factory, its production inference platform." NVIDIA adds that "purpose-built AI inference cloud Groq plans to be among the platform's earliest adopters" as well.
Context
Groq 3 LPX is the first shipping product to come out of NVIDIA's roughly $20 billion acquisition of chip startup Groq's assets, announced late last year. Groq built its name on LPU (language processing unit) architecture optimized specifically for low-latency token generation, in contrast to general-purpose GPU training silicon. Folding that design into the Vera Rubin platform gives NVIDIA a dedicated inference-side complement to the GPUs it sells for training and batch workloads.
The announcement lands at Hot Chips 2026, the annual conference where NVIDIA, AMD, and other silicon makers typically detail forthcoming architectures. It follows NVIDIA's Vera Rubin NVL72 systems ramping into production earlier this month across CoreWeave, Google Cloud, and Nebius — Groq 3 LPX is a distinct, inference-specialized product line rather than a variant of that general-purpose rollout.
Why it matters
Inference cost and latency, not just training throughput, have become the binding constraint for agentic AI products — coding assistants, voice agents, and other interactive systems that need to stream tokens back with minimal delay. A dedicated accelerator claiming a 4x responsiveness edge over alternatives, if it holds up under independent load, gives NVIDIA an answer to specialized inference chips from Groq's former rivals (Cerebras, SambaNova) and cloud providers building their own silicon.
Nebius's early adoption also signals that neoclouds — the newer generation of AI-focused cloud providers competing for inference workloads against the hyperscalers — see enough differentiation in Groq 3 LPX to commit ahead of broader general availability. Whether NVIDIA can scale production fast enough to meet agentic-AI demand, and whether the token-per-second gains translate into meaningfully lower cost per query for customers, are the questions that will determine how much this reshapes the inference market over the next year.
Corroborating sources
- Nvidianews.nvidia
https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-full-production-with-world-class-speed-for-agentic-ai
“NVIDIA Groq 3 LPX, the interactive AI inference accelerator, is now in full production.”
- Cnbc
https://www.cnbc.com/2026/08/24/nvidia-says-groq-racks-will-be-online-this-year-after-20-billion-deal.html