With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA has begun full production of the Groq 3 LPX accelerator, designed to enhance token generation speed for agentic AI systems within the Vera Rubin NVL72 platform. Early adopters include Nebius, CoreWeave, and SpaceXAI, signaling a shift toward optimized inference infrastructure for real-time AI applications.
NVIDIA announced that its Groq 3 LPX accelerator has entered full production, integrated into the Vera Rubin NVL72 platform to support agentic AI systems. In benchmark tests using the open-source Gemma 4 31B model, the system achieved 3,400 output tokens per second for long-context workloads, four times faster than competing platforms. The announcement underscores NVIDIA’s focus on optimizing inference performance as AI transitions from training to reasoning and multi-agent collaboration.
Industry partners are adopting the Vera Rubin platform to power next-generation AI workloads. SpaceXAI will use NVIDIA Vera CPUs to accelerate agentic AI tasks such as orchestration and code execution, while CoreWeave has deployed Spectrum-X Multiplane to connect Vera Rubin racks with high-bandwidth, lossless networks. Nebius is the first AI cloud to integrate Groq 3 LPX, enabling developers to build highly responsive agentic applications at scale.
The shift toward agentic AI demands infrastructure optimized for low-latency token generation and large-scale context processing. NVIDIA Groq 3 LPX is designed to address decode latency challenges by accelerating token generation alongside Vera Rubin GPUs, which handle context processing. The combined architecture aims to eliminate trade-offs between speed and throughput, improving responsiveness and infrastructure efficiency for agentic workloads.
NVIDIA describes the evolving infrastructure as a 'token factory,' emphasizing performance, throughput, and economic efficiency. Groq 3 LPX extends the Vera Rubin NVL72 platform with 256 LP30 accelerators per rack, connected via chip-to-chip links to form a unified inference engine. This design supports deterministic, high-speed inference for modern AI factories, enabling real-time agent interactions and scalable deployment across data centers and orbital systems.