OFICIAL NVIDIA Newsroom

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents

What happened
Based on NVIDIA Newsroom · Aug 24, 2026

NVIDIA has begun full production of the Groq 3 LPX accelerator, designed to enhance token generation speed for agentic AI systems within the Vera Rubin NVL72 platform. Early adopters include Nebius, CoreWeave, and SpaceXAI, signaling a shift toward optimized inference infrastructure for real-time AI applications.

With Groq 3 LPX in Full Production, NVIDIA Extends Vera Rubin Inference for Agents
NVIDIA Newsroom — NVIDIA
Key points
·
The next era of AI inference won’t be defined by a single breakthrough chip, network or system.
·
It’ll be defined by how every layer of the AI factory works together.
·
That’s why NVIDIA is extending Vera Rubin NVL72 with fast token generation for agentic systems.
·
NVIDIA details today, the NVIDIA Vera Rubin rack-scale system NVIDIA Groq 3 LPX is in full production.
Key numbers
·
In benchmark tests using the open-source Gemma 4 31B model, the system achieved 3,400 output tokens per second for long-context workloads, four times faster than competing platforms.

NVIDIA announced that its Groq 3 LPX accelerator has entered full production, integrated into the Vera Rubin NVL72 platform to support agentic AI systems. In benchmark tests using the open-source Gemma 4 31B model, the system achieved 3,400 output tokens per second for long-context workloads, four times faster than competing platforms. The announcement underscores NVIDIA’s focus on optimizing inference performance as AI transitions from training to reasoning and multi-agent collaboration.

Industry partners are adopting the Vera Rubin platform to power next-generation AI workloads. SpaceXAI will use NVIDIA Vera CPUs to accelerate agentic AI tasks such as orchestration and code execution, while CoreWeave has deployed Spectrum-X Multiplane to connect Vera Rubin racks with high-bandwidth, lossless networks. Nebius is the first AI cloud to integrate Groq 3 LPX, enabling developers to build highly responsive agentic applications at scale.

The shift toward agentic AI demands infrastructure optimized for low-latency token generation and large-scale context processing. NVIDIA Groq 3 LPX is designed to address decode latency challenges by accelerating token generation alongside Vera Rubin GPUs, which handle context processing. The combined architecture aims to eliminate trade-offs between speed and throughput, improving responsiveness and infrastructure efficiency for agentic workloads.

NVIDIA describes the evolving infrastructure as a 'token factory,' emphasizing performance, throughput, and economic efficiency. Groq 3 LPX extends the Vera Rubin NVL72 platform with 256 LP30 accelerators per rack, connected via chip-to-chip links to form a unified inference engine. This design supports deterministic, high-speed inference for modern AI factories, enabling real-time agent interactions and scalable deployment across data centers and orbital systems.

Original source → Deals on Clipraptor.com →