From Training to Production, NVIDIA and CoreWeave Close the Loop on Agentic AI
NVIDIA and CoreWeave announced production availability of Vera Rubin NVL72 systems and the NVIDIA Vera CPU, with Cognition deploying Devin AI workloads on CoreWeave Cloud.
CoreWeave introduced NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet networking into production at its Fully Connected event in San Francisco. The company also launched CoreWeave Forge, a unified environment for training, evaluating and improving AI models and agents on NVIDIA accelerated computing. Cognition, the developer of the Devin AI software engineer, is the first customer running production workloads on Vera Rubin, scaling to thousands of GPUs in nine months.
Cognition benchmarked Vera Rubin NVL72 against a GB200 NVL72 baseline using real-world software engineering tasks from FrontierCode. Early tests showed Vera Rubin delivered up to a 4.8x increase in total token throughput for SWE-2 inference workloads, enabling faster real-time code generation and more responsive multistep reasoning for Devin. CoreWeave achieved these results through close collaboration with NVIDIA, standing up a production Vera Rubin cluster for Cognition in days.
CoreWeave also announced availability of the NVIDIA Vera CPU, purpose-built for agentic AI workloads. In a single rack, CoreWeave’s deployment of Vera provides 128 CPUs and 11,264 cores, supporting more than 11,000 concurrent isolated agent environments. CoreWeave Sandboxes, powered by Vera CPUs and BlueField-4 DPUs, enable secure, high-performance agent communication at low latency, with startup times more than 3x faster than previous configurations.
CoreWeave Forge integrates tools like Weights & Biases, OpenPipe and the open source marimo notebook project into a single environment for continuous model and agent improvement. It supports reinforcement learning without redeploys through RL Rollouts, which loads new checkpoints into live deployments. Early customers include Canva, Capital One and MasterClass, while Ennoble Care selected CoreWeave to run clinical AI inference using reserved NVIDIA RTX PRO 6000 GPU capacity.