VFF - The signal in the noise
NewsTrending

NVIDIA Blackwell Leads First Agentic AI Benchmark

Read original
Share
NVIDIA Blackwell Leads First Agentic AI Benchmark

Artificial Analysis released AgentPerf, the first benchmark designed specifically for agentic AI workloads, showing NVIDIA's Blackwell Ultra NVL72 platform delivering 20x more agents per megawatt than Hopper-based systems. The benchmark reflects the fundamentally different performance characteristics of agentic AI, which chains dozens to hundreds of LLM calls with tool execution rather than single-turn completions. Results are based on real coding agent trajectories across 12+ programming languages, providing infrastructure providers and enterprises with direct metrics for deployment decisions.

  • AgentPerf is the first benchmark built specifically for agentic AI, measuring concurrent agent capacity and responsiveness rather than single LLM call speed
  • NVIDIA GB300 NVL72 runs up to 20x more agents per megawatt than HGX H200 systems on DeepSeek V4 Pro workloads
  • Agentic AI differs fundamentally from conversational AI: agents chain dozens to hundreds of LLM calls with tool calls, creating multiplicative complexity rather than additive
  • Benchmark methodology uses real coding agent trajectories from public repositories, with tool calls simulated to isolate accelerated computing performance

Existing AI inference benchmarks measure single LLM calls and were not designed for agentic workloads, where chained calls, tool delays, and growing context create fundamentally different performance stresses. AgentPerf fills this gap by measuring what actually matters for production agentic AI: concurrent agent capacity and responsiveness at scale. This enables infrastructure providers and enterprises to make informed deployment decisions based on real-world agentic patterns.

For enterprises deploying AI agents at scale, infrastructure efficiency directly impacts cost per concurrent agent and power consumption. AgentPerf translates benchmark results into actionable metrics: how many concurrent agentic tasks can run per accelerator and per megawatt of power. NVIDIA's 20x advantage on this benchmark could significantly influence infrastructure purchasing decisions for agentic AI deployments.

  • Agentic AI performance cannot be accurately assessed using conversational AI benchmarks, creating demand for specialized measurement tools and potentially invalidating prior infrastructure comparisons
  • NVIDIA's Blackwell architecture appears optimized for agentic workloads through rack-scale GPU coordination, CUDA kernel optimization for expert distribution, and TensorRT LLM efficiency gains
  • Infrastructure decisions for agentic AI deployments will increasingly be based on concurrent agent capacity and power efficiency rather than raw inference speed metrics

Monitor whether other accelerator providers publish AgentPerf results and how their performance compares to NVIDIA's baseline. Watch for adoption of AgentPerf as an industry standard for agentic AI infrastructure evaluation. Track whether the 20x efficiency advantage translates into actual market share gains for Blackwell in agentic AI deployments.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Arizona chip boom faces water shortage threat

Arizona chip boom faces water shortage threat

Arizona is set to lose more than a quarter of its annual Colorado River water allocation due to a recent federal water management decision. The state, home to major semiconductor manufacturing expansions by TSMC and Intel, depends on the Colorado River for over a third of its water supply. The reduction comes as the river experiences unprecedented drought, creating a significant constraint on water-intensive chip production in a region critical to U.S. semiconductor revival efforts.

by Justine Calma· The Verge AI
Nvidia projects 70% growth, denies circular dealing

Nvidia projects 70% growth, denies circular dealing

Nvidia CEO Jensen Huang stated the company expects to grow 70% in the coming year, citing its broad involvement across multiple business segments. Huang addressed concerns about circular dealing, asserting that Nvidia's various business relationships are not self-referential. The statement reflects confidence in sustained demand across the company's portfolio.

by Julie Bort· TechCrunch AI
Microsoft to Triple Azure Capacity by 2032 Amid Server Shortage

Microsoft to Triple Azure Capacity by 2032 Amid Server Shortage

Microsoft plans to triple Azure's data center capacity to over 38 gigawatts by 2032, up from 12 gigawatts currently. The expansion reflects the company's response to server shortages constraining its cloud operations. The buildout will add approximately 26 gigawatts of new compute capacity over the next six years.

by Aaron Holmes· The Information
Skild AI's S1 Robot Learns New Tasks From Single Video
TrendingNews

Skild AI's S1 Robot Learns New Tasks From Single Video

Skild AI launched S1, a robot foundation model that learns new tasks from single video demonstrations without retraining, using in-context learning to adapt to dynamic manufacturing and warehouse environments. Built on NVIDIA infrastructure, S1 achieved a 66% success rate per step on unfamiliar multistep tasks, compared to 9% for competing systems. The company has reached $100 million annual revenue run rate within 10 months of first commercial deployment, with over 60 partnerships across manufacturing, logistics, and other sectors.

by Sasa Docca· NVIDIA Blog (AI)