VFF - The signal in the noise
News

NVIDIA Vera Rubin Cuts Agentic AI Costs 35x, Boosts Efficiency 30x

Read original
Share
NVIDIA Vera Rubin Cuts Agentic AI Costs 35x, Boosts Efficiency 30x

NVIDIA's Vera Rubin NVL72 GPU system delivers up to 30x higher throughput per megawatt than its GB300 NVL72 predecessor on agentic AI workloads, according to measurements using the SemiAnalysis AgentX benchmark. Agentic tasks consume 15x more tokens than simple chat because agents iteratively query databases, invoke sub-agents, and accumulate context across multiple reasoning steps. For power-constrained AI infrastructure operators, the efficiency gain translates directly to running significantly more agent-based work within the same energy budget.

  • Vera Rubin NVL72 achieves 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads, measured using SemiAnalysis AgentX benchmark with real-world coding sessions
  • Token cost per million drops by up to 35x on Vera Rubin versus GB300 NVL72, enabling continuous agent operation at scale
  • Agentic AI workloads consume 15x more tokens than chat because agents iteratively reason, call tools, spawn sub-agents, and accumulate context across steps
  • NVIDIA DSX MaxLPS power management can provision up to 40% more GPUs within the same megawatt budget, further improving throughput efficiency

Agentic AI is moving into production across industries, but these systems generate vastly higher token consumption than traditional chat interfaces due to iterative reasoning and tool calling. Infrastructure efficiency directly determines whether AI factories can run agents profitably. Vera Rubin's efficiency gains address a critical bottleneck for operators managing power-constrained data centers.

For AI infrastructure operators and service providers, throughput per megawatt determines revenue while cost per token determines profit margin. A 30x efficiency improvement on agentic workloads allows operators to either serve 30x more agent requests within existing power budgets or reduce operational costs significantly. This directly impacts the unit economics of agentic AI services.

  • Agentic AI workloads require fundamentally different performance measurement approaches than chat, since context can reach hundreds of thousands of tokens with high variability in input and output lengths
  • Power efficiency has become a primary competitive factor in GPU design for inference, not just raw throughput, as agentic AI scales into production
  • Long-context handling and tool-calling optimization are now central to GPU architecture decisions, reflected in Vera Rubin's codesign approach

Monitor whether Vera Rubin's efficiency gains hold across diverse agentic models and use cases beyond the tested set (Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, DeepSeek V4 Pro). Watch for SemiAnalysis's formal review of these results and whether competing GPU makers (AMD, Intel, others) publish comparable agentic workload benchmarks. Track whether the 35x token cost reduction translates into lower pricing for agentic AI services in the market.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

NVIDIA and Microsoft Launch RTX Spark for Local AI on Windows
TrendingNews

NVIDIA and Microsoft Launch RTX Spark for Local AI on Windows

NVIDIA and Microsoft announced RTX Spark, a new AI hardware and software platform designed to run AI agents locally on Windows PCs. RTX Spark combines NVIDIA's Blackwell GPU with Grace CPU, offering up to 128GB unified memory and one petaflop of FP4 AI performance. Laptop preorders begin today with availability on October 16, while compact desktops launch in November. Microsoft also released Microsoft Execution Containers (MXC) as OS-level infrastructure to run agents securely in the background.

by Gerardo Delgado· NVIDIA Blog (AI)
Ex-Google, Nvidia Execs Launch GPU Access Alternative

Ex-Google, Nvidia Execs Launch GPU Access Alternative

Former executives from Google, Nvidia, and Apple, along with ex-Andreessen Horowitz partner Anjney Midha, have launched a new company aimed at reducing GPU access barriers for smaller companies and startups. The venture addresses a growing compute crunch as major cloud providers like Microsoft tighten control over GPU availability. The move signals growing demand for alternative pathways to affordable AI infrastructure outside dominant cloud platforms.

by Phoebe Liu· The Information
SpaceX Seeks $40B to Buy Nvidia Chips in Apollo-Led Round
TrendingNews

SpaceX Seeks $40B to Buy Nvidia Chips in Apollo-Led Round

SpaceX is seeking to raise $40 billion in a financing round led by Apollo Global Management, with proceeds earmarked for purchasing Nvidia chips. The deal, reported by the Financial Times, is expected to close in 2027 and will consist of approximately $10 billion in bank loans and $30 billion in other financing. The capital raise underscores SpaceX's significant infrastructure needs as it expands AI and computing capabilities alongside its space operations.

by Tiffany Li· The Information
Lambda raises $4B for 2027 IPO as AI infrastructure matures
TrendingNews

Lambda raises $4B for 2027 IPO as AI infrastructure matures

Lambda, an AI computing startup backed by Nvidia, is raising up to $4 billion at a $14.5 billion pre-money valuation ahead of a planned 2027 IPO. The funding round is led by Coatue and Blackstone. The raise positions Lambda for public markets entry as demand for AI infrastructure continues to grow.

by Rebecca Bellan· TechCrunch AI