Vera Rubin NVL72 Debuts With 3.7x Throughput Gain Over GB300
NVIDIA's Vera Rubin NVL72 system achieved leading performance in its MLPerf Inference v6.1 debut, delivering up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. The results demonstrate the effectiveness of full-stack hardware and software codesign, including enhanced Tensor Cores, NVFP4 precision, and disaggregated serving techniques. A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, and software optimizations alone delivered up to 1.6x performance gains from v6.0 to v6.1.
TL;DR
- Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL, up to 2.5x on DeepSeek-R1
- 288-GPU multi-rack submission achieved 99% scaling efficiency with nearly linear throughput growth
- Software optimizations between v6.0 and v6.1 alone drove up to 1.6x performance gains
- Results reflect full-stack codesign including enhanced Tensor Cores, NVFP4 precision, and disaggregated serving with expert parallelism
Why It Matters
MLPerf Inference benchmarks are industry-standard measures of AI inference performance. Vera Rubin's strong debut results signal meaningful progress in inference efficiency and throughput, which directly impact the economics of serving large language models at scale. The 99% scaling efficiency across multiple racks demonstrates that performance gains translate reliably as infrastructure scales.
Business Impact
Higher inference throughput means more tokens generated per rack, more concurrent users served, and lower cost per token. Organizations evaluating AI infrastructure investments use MLPerf results to assess long-term inference economics. Vera Rubin's performance advantage could influence purchasing decisions for large-scale inference deployments.
Key Implications
- Vera Rubin NVL72 offers significantly better token generation economics than GB300 NVL72, potentially shifting infrastructure ROI calculations for inference-heavy workloads
- 99% scaling efficiency suggests NVIDIA's interconnect and software stack can reliably scale inference across multiple racks without proportional performance loss
- Continued software optimization post-submission indicates further performance gains are likely, making current benchmarks a floor rather than a ceiling for Vera Rubin performance
What to Watch
Monitor whether Vera Rubin's performance advantage persists across additional MLPerf benchmarks and real-world inference workloads beyond the two tested models. Track adoption rates among cloud providers and enterprises making infrastructure decisions. Watch for competing systems from other vendors and whether software optimizations continue to drive gains in future MLPerf rounds.
Subscribe to the newsletter
The latest stories and analysis, delivered to your inbox.
Free. No spam. Unsubscribe any time.



