Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

CoreWeave opens Vera Rubin NVL72 cloud access as Cognition runs production workloads

CoreWeave made NVIDIA Vera Rubin NVL72 available to select cloud customers, with Cognition running production workloads on a new cluster.

D
Sep 30, 2026 · 2 min read

CoreWeave made NVIDIA Vera Rubin NVL72 systems available to select cloud customers and deployed a production cluster for Cognition, NVIDIA said on September 30. NVIDIA and CoreWeave named Cognition, the company behind the Devin software-engineering agent, as the first customer running production workloads on the platform through CoreWeave Cloud.

CoreWeave described the service as limited availability, with hundreds of Rubin GPUs deployed across multiple regions. According to CoreWeave, Cognition engineers began running production workloads within days of the racks being handed over. The company did not disclose pricing, available capacity, named regions or a date for general availability.

NVIDIA said Cognition uses CoreWeave infrastructure for Devin training, reinforcement learning and production inference. For an early comparison with the previous-generation GB200 NVL72, Cognition sampled software-engineering tasks from FrontierCode and deployed agents to solve them. Cognition said Vera Rubin delivered up to 4.8 times as much total token throughput for SWE-2 inference. CoreWeave said the comparison was per GPU at matched interactivity, meaning the systems were tested while holding per-user responsiveness constant.

CoreWeave separately reported that Vera Rubin produced 3.8 times the output-token throughput per GPU of GB200 NVL72 for reinforcement-learning workloads at matched interactivity. NVIDIA said the inference result should translate into faster real-time code generation and more responsive multistep reasoning for Devin. The companies did not publish raw benchmark data, a reproducible test configuration or an independent audit.

According to CoreWeave, each Vera Rubin NVL72 rack combines 72 Rubin GPUs and 36 Vera CPUs with NVLink 6, ConnectX-9 SuperNICs and BlueField-4 data-processing units. The rack-scale design connects the GPUs so they can operate as one system for training and inference. On September 16, CoreWeave said it had connected multiple racks, creating a scale-out cluster spanning hundreds of Rubin GPUs.

The Cognition measurements are separate from an earlier CoreWeave-run efficiency test. In July, CoreWeave said Vera Rubin generated 10 times more tokens per second per megawatt than GB200 NVL72 at matched interactivity while running DeepSeek R1. That company-reported test did not use Cognition’s SWE-2 workload.

More news