VFF - The signal in the noise
Research

Vera Rubin NVL72 Debuts With 3.7x Throughput Gain Over GB300

Read original
Share
Vera Rubin NVL72 Debuts With 3.7x Throughput Gain Over GB300

NVIDIA's Vera Rubin NVL72 system achieved leading performance in its MLPerf Inference v6.1 debut, delivering up to 3.7x higher throughput than the GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. The results demonstrate the effectiveness of full-stack hardware and software codesign, including enhanced Tensor Cores, NVFP4 precision, and disaggregated serving techniques. A 288-GPU submission across four GB300 NVL72 racks achieved 99% scaling efficiency, and software optimizations alone delivered up to 1.6x performance gains from v6.0 to v6.1.

  • Vera Rubin NVL72 delivers up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL, up to 2.5x on DeepSeek-R1
  • 288-GPU multi-rack submission achieved 99% scaling efficiency with nearly linear throughput growth
  • Software optimizations between v6.0 and v6.1 alone drove up to 1.6x performance gains
  • Results reflect full-stack codesign including enhanced Tensor Cores, NVFP4 precision, and disaggregated serving with expert parallelism

MLPerf Inference benchmarks are industry-standard measures of AI inference performance. Vera Rubin's strong debut results signal meaningful progress in inference efficiency and throughput, which directly impact the economics of serving large language models at scale. The 99% scaling efficiency across multiple racks demonstrates that performance gains translate reliably as infrastructure scales.

Higher inference throughput means more tokens generated per rack, more concurrent users served, and lower cost per token. Organizations evaluating AI infrastructure investments use MLPerf results to assess long-term inference economics. Vera Rubin's performance advantage could influence purchasing decisions for large-scale inference deployments.

  • Vera Rubin NVL72 offers significantly better token generation economics than GB300 NVL72, potentially shifting infrastructure ROI calculations for inference-heavy workloads
  • 99% scaling efficiency suggests NVIDIA's interconnect and software stack can reliably scale inference across multiple racks without proportional performance loss
  • Continued software optimization post-submission indicates further performance gains are likely, making current benchmarks a floor rather than a ceiling for Vera Rubin performance

Monitor whether Vera Rubin's performance advantage persists across additional MLPerf benchmarks and real-world inference workloads beyond the two tested models. Track adoption rates among cloud providers and enterprises making infrastructure decisions. Watch for competing systems from other vendors and whether software optimizations continue to drive gains in future MLPerf rounds.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Lightweight dual-model agents show promise for autonomous materials research
Research

Lightweight dual-model agents show promise for autonomous materials research

Researchers at Nature Machine Intelligence have demonstrated a dual-model architecture for autonomous crystal materials research using two lightweight large language models working collaboratively. The approach combines reasoning and scientific tool execution while maintaining computational efficiency and local deployability. The method achieves competitive performance without requiring expensive infrastructure, making advanced materials research more accessible.

by Tongyu Shi· Nature Machine Intelligence
Perplexity deploys GPT-6 Astra to autonomous production systems
News

Perplexity deploys GPT-6 Astra to autonomous production systems

Perplexity is deploying OpenAI's GPT-6 Astra model to handle end-to-end system operations, including writing communications, modifying software, and monitoring production infrastructure. The company reports significantly reduced oversight requirements compared to earlier models. This represents a shift toward autonomous AI management of critical business systems.

· OpenAI
Alibaba Open-Sources Qwen3.8 Trillion-Parameter Model
News

Alibaba Open-Sources Qwen3.8 Trillion-Parameter Model

Alibaba released Qwen3.8-2.4T-A95B as open weights on August 12, 2026, marking the first time a Qwen-Max-class model became publicly available. The 2.4 trillion parameter model uses a hybrid linear-plus-full-attention architecture with 95 billion activated parameters per token and supports up to 262K native context tokens, extensible to 1M. AWS published a deployment guide showing how to run the model on SageMaker HyperPod using vLLM on ml.p6-b300 instances with NVIDIA B300 Blackwell Ultra GPUs.

by Dmitry Soldatkin· AWS Machine Learning Blog
Saudi Arabia Launches Arabic AI Model With Chinese Partner
TrendingNews

Saudi Arabia Launches Arabic AI Model With Chinese Partner

Humain, Saudi Arabia's state-owned AI company, announced the humain-m3 model, an Arabic language model built on Chinese firm MiniMax's open-source M3 foundation. The model was pre-trained on more than 1 trillion tokens of Arabic content. The development represents a collaboration between Saudi and Chinese AI capabilities focused on Arabic language processing.

by Juro Osawa· The Information