AI Infrastructure
NVIDIA puts Groq 3 LPX into production as AI agents push inference beyond GPUs alone
NVIDIA said Groq 3 LPX is now in full production, extending Vera Rubin systems with faster token generation for long-context and agentic AI workloads.
NVIDIA has moved Groq 3 LPX into full production, using the Hot Chips conference to frame inference as the next major constraint in the AI factory. The company said on August 24 that the interactive inference accelerator extends its Vera Rubin NVL72 platform by dramatically increasing token generation speed for agentic systems. NVIDIA’s message is that the next phase of AI infrastructure will not be defined only by faster training, but by how quickly deployed models can read long context, reason through multiple steps and respond without making users wait.
Groq 3 LPX is designed around a specific pain point in modern inference. AI agents often work through hundreds or thousands of steps: reading files, calling tools, checking outputs, writing code, revising plans and repeating the loop. Those steps create huge volumes of tokens, especially when the model has to keep a long context window active. GPUs remain powerful for many parts of that workload, but NVIDIA is arguing that the generation phase needs specialized acceleration when responsiveness and cost per token become the commercial bottleneck.
In benchmark results cited by NVIDIA, Groq 3 LPX delivered about 3,400 output tokens per second while running the open-source Gemma 4 31B model with a 100,000-token context, which the company described as the fastest recorded performance for that model. NVIDIA also said the system provided roughly four times faster responsiveness than the nearest alternative platform for latency-sensitive workloads. Vendor benchmarks always need scrutiny, but the numbers illustrate why inference is becoming a separate design problem from training. A model that is strong on paper may feel slow or expensive if it cannot generate and verify steps quickly enough in production.
The product also shows NVIDIA’s strategy after licensing inference technology from Groq and hiring key talent connected to that company. Groq 3 LPX is meant to work alongside Vera Rubin racks rather than replace GPUs outright. The architecture reflects a broader industry idea known as disaggregation, where different phases of model serving are assigned to processors optimized for those jobs. Long-context processing, token generation, networking and storage can be engineered as one system rather than treated as independent components bolted together inside a data center.
Nebius is the first AI cloud provider named by NVIDIA as an adopter, with plans to bring Groq 3 LPX into its Token Factory inference platform. Groq, the purpose-built inference cloud company whose technology underpins part of the system, is also listed among the earliest planned adopters. That matters because cloud availability is what turns a chip announcement into something developers can actually test. If enterprises can access faster agent inference through familiar APIs, the impact may show up first in coding agents, research assistants, customer support automation and workflow tools where delays compound across many model calls.
The announcement lands as AI spending is under closer scrutiny. Companies have poured capital into training clusters and data centers, but many practical AI workloads now run continuously after deployment. Inference costs can accumulate every time a user asks a question or an agent takes a step. Lower latency and lower token cost therefore become business questions, not just engineering achievements. NVIDIA is trying to show that it can dominate that production phase just as it dominated the training boom.
Groq 3 LPX will still need real-world validation outside controlled benchmarks, and customers will weigh performance against price, availability, software maturity and energy use. Yet the direction is clear. As AI agents become less like chatbots and more like systems that work over long horizons, infrastructure must support rapid, repeated reasoning. NVIDIA’s new production milestone is a signal that the AI hardware race is moving from building bigger brains to making deployed intelligence fast enough to use all day.