VFF - The signal in the noise
News

NVIDIA Vera Rubin Cuts Agentic AI Costs 35x, Boosts Efficiency 30x

Read original
Share
NVIDIA Vera Rubin Cuts Agentic AI Costs 35x, Boosts Efficiency 30x

NVIDIA's Vera Rubin NVL72 GPU system delivers up to 30x higher throughput per megawatt than its GB300 NVL72 predecessor on agentic AI workloads, according to measurements using the SemiAnalysis AgentX benchmark. Agentic tasks consume 15x more tokens than simple chat because agents iteratively query databases, invoke sub-agents, and accumulate context across multiple reasoning steps. For power-constrained AI infrastructure operators, the efficiency gain translates directly to running significantly more agent-based work within the same energy budget.

  • Vera Rubin NVL72 achieves 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads, measured using SemiAnalysis AgentX benchmark with real-world coding sessions
  • Token cost per million drops by up to 35x on Vera Rubin versus GB300 NVL72, enabling continuous agent operation at scale
  • Agentic AI workloads consume 15x more tokens than chat because agents iteratively reason, call tools, spawn sub-agents, and accumulate context across steps
  • NVIDIA DSX MaxLPS power management can provision up to 40% more GPUs within the same megawatt budget, further improving throughput efficiency

Agentic AI is moving into production across industries, but these systems generate vastly higher token consumption than traditional chat interfaces due to iterative reasoning and tool calling. Infrastructure efficiency directly determines whether AI factories can run agents profitably. Vera Rubin's efficiency gains address a critical bottleneck for operators managing power-constrained data centers.

For AI infrastructure operators and service providers, throughput per megawatt determines revenue while cost per token determines profit margin. A 30x efficiency improvement on agentic workloads allows operators to either serve 30x more agent requests within existing power budgets or reduce operational costs significantly. This directly impacts the unit economics of agentic AI services.

  • Agentic AI workloads require fundamentally different performance measurement approaches than chat, since context can reach hundreds of thousands of tokens with high variability in input and output lengths
  • Power efficiency has become a primary competitive factor in GPU design for inference, not just raw throughput, as agentic AI scales into production
  • Long-context handling and tool-calling optimization are now central to GPU architecture decisions, reflected in Vera Rubin's codesign approach

Monitor whether Vera Rubin's efficiency gains hold across diverse agentic models and use cases beyond the tested set (Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, DeepSeek V4 Pro). Watch for SemiAnalysis's formal review of these results and whether competing GPU makers (AMD, Intel, others) publish comparable agentic workload benchmarks. Track whether the 35x token cost reduction translates into lower pricing for agentic AI services in the market.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Nvidia Lands SpaceX, Nebius as Early Vera CPU Customers
TrendingNews

Nvidia Lands SpaceX, Nebius as Early Vera CPU Customers

Nvidia announced that SpaceX and AI cloud company Nebius will be early customers for its Vera CPU and Groq LPX inference-focused racks, marking the company's expansion beyond its core GPU business. The move addresses investor and analyst questions about how Nvidia plans to diversify its product portfolio beyond graphics processors. Both products target different segments of the AI infrastructure market, with Vera serving as a central processing unit option and Groq LPX designed for fast inference workloads.

by Phoebe Liu· The Information
IBM mainframe chip runs Arm and Z workloads on same cores
TrendingNews

IBM mainframe chip runs Arm and Z workloads on same cores

IBM announced a dual-architecture mainframe processor at Hot Chips that can natively execute both Arm and IBM Z instruction sets on the same cores, switching between them in nanoseconds. Built on 2-nanometer process technology with 11 cores running above 5.7 GHz, the chip allows enterprises to run Arm-native Linux software and AI frameworks alongside z/OS transaction-processing workloads on shared silicon. The design represents the first hardware outcome of IBM and Arm's April strategic collaboration and directly addresses whether mainframes can remain relevant in an AI-dominated infrastructure landscape.

by michael.nunez@venturebeat.com (Michael Nuñez)· VentureBeat AI
General Intuition raises $6B at valuation for embodied AI
TrendingNews

General Intuition raises $6B at valuation for embodied AI

General Intuition, an AI startup developing foundation models for generalized agents that navigate physical and temporal space, is raising capital at a $6 billion pre-money valuation from investors including Valor Ventures, Point72 Ventures, and Seven Seven Six. The funding round signals investor confidence in AI systems designed for robotics and embodied AI applications. The startup's focus on spatial reasoning and agent movement represents a shift toward practical, physical-world AI deployment.

by Rebecca Bellan· TechCrunch AI
Nvidia Raises Flagship AI Chip Prices 17%

Nvidia Raises Flagship AI Chip Prices 17%

Nvidia is raising prices for its Grace Blackwell and Vera Rubin flagship AI chip systems by approximately 17%, according to server makers. The increase adds to mounting cost pressures facing cloud providers and data center developers, who are already contending with power infrastructure delays and other unexpected expenses. The price hike reflects Nvidia's market position as demand for advanced AI chips remains strong despite the growing cost burden on customers.

by Amir Efrati· The Information