Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

NVIDIA Vera Rubin NVL72 debuts in MLPerf Inference v6.1 as a Preview system

MLCommons' first reviewed Vera Rubin results show higher submitted throughput than a same-size GB300 NVL72, while leaving power and economic claims untested.

D
Sep 18, 2026 · 2 min read

NVIDIA has submitted Vera Rubin NVL72 to MLPerf Inference v6.1, giving infrastructure buyers the first reviewed and audited benchmark results for the rack-scale platform. MLCommons identified Vera Rubin as a new Preview system and said submissions using it produced the round’s largest per-accelerator gains on the vision-language and DeepSeek-R1 workloads.

The published results ledger lists the system as NVIDIA VR200 NVL72, with 72 VR200 accelerators and NVIDIA Vera CPUs. Its six Closed-division Datacenter results cover Interactive, Offline and Server scenarios for both DeepSeek-R1 and Qwen3-VL-235B-A22B. That classification matters: MLCommons lists the Vera Rubin configuration as Preview, while the 72-accelerator GB300 NVL72 systems used for comparison are Available.

On DeepSeek-R1, the Vera Rubin submission recorded 652,750.14 tokens per second in Interactive, 1,183,326.85 in Offline and 1,175,890.22 in Server. The corresponding GB300 NVL72 results were 253,506, 679,740 and 596,944 tokens per second. Across the three scenarios, the submitted figures work out to gains of about 1.74 times to 2.57 times.

On Qwen3-VL, Vera Rubin recorded 1,306.61 queries per second in Interactive, 2,392.70 samples per second in Offline and 2,323.27 queries per second in Server. GB300 NVL72 recorded 349.29, 1,304.99 and 1,210.49 in the same scenarios and units. The submission-to-submission ratios range from about 1.83 times to 3.74 times. They align with NVIDIA’s rounded claims of up to 2.5 times for DeepSeek-R1 and 3.7 times for Qwen3-VL.

MLPerf’s Closed division holds the model mathematically equivalent to a reference implementation, supporting comparisons between submitted hardware and software configurations. It does not isolate the chip’s effect. NVIDIA said its Qwen3-VL entry used vLLM with NVIDIA Dynamo, while its DeepSeek-R1 entry used TensorRT-LLM; the company also cited disaggregated serving and expert parallelism.

Preview is an MLPerf availability category, not a claim that no hardware exists. Under the benchmark’s availability rules, Available systems contain components offered for purchase or cloud rental, while Preview systems must be submitted as Available in the next round. For context on that distinction, DataPhoenix previously covered NVIDIA’s Vera CPU shipping update. NVIDIA said separately in July that production was ramping and racks were running at several partners, but the exact benchmark configuration remains Preview in version 6.1.

NVIDIA said the increased throughput could support more users, more revenue and lower cost per token. The six Vera Rubin result rows do not include full-system power measurements, however, and MLCommons validates power only when those measurements accompany a submission. The version 6.1 evidence therefore establishes throughput for the submitted configurations, not performance per watt, revenue, total cost of ownership or cost per token.

More news