VFF - The signal in the noise
News

NVIDIA Claims 5x Token Cost Cuts on Blackwell via Software Stack

Read original
Share
NVIDIA Claims 5x Token Cost Cuts on Blackwell via Software Stack

NVIDIA claims its inference software stack has reduced token costs by up to 5x on the DeepSeek V4 model within one month on its Blackwell platform. The company argues that as AI moves from pilots to production, software optimization across serving, acceleration, and infrastructure layers becomes critical to cost efficiency. Leading inference providers including Baseten, Cognition, Deep Infra, and Together AI is already deploying these tools to improve throughput and reduce latency on Blackwell GPUs.

  • NVIDIA reports 5x token cost reduction on DeepSeek V4 in one month using Blackwell platform and its inference software stack
  • Baseten achieved 50% more tokens per second using TensorRT-LLM on Blackwell for DeepSeek V4 Pro
  • DigitalOcean and Hippocratic AI increased inference throughput 30% while maintaining sub-half-second response time across 10 million patient calls
  • NVIDIA's three-layer software approach (production operation, application acceleration, infrastructure access) aims to turn distributed agentic AI workloads into lower-cost serving

As AI workloads shift from experimental pilots to production factories, cost per token has become the primary infrastructure metric, replacing peak performance specs. NVIDIA's software stack claims to compound optimizations across hardware, networking, and serving layers, directly impacting the economics of running large language models at scale. Early adopters report significant throughput gains and cost reductions, suggesting software optimization may become as important as hardware selection.

Organizations deploying AI in production face pressure to reduce inference costs while meeting latency requirements. NVIDIA's software stack enables inference providers to serve models more efficiently without building custom infrastructure from scratch, lowering barriers to competitive deployment. Companies like Cognition and Cursor can accelerate time-to-production by using ready-made frameworks rather than building serving infrastructure independently.

  • Software optimization on existing hardware may deliver cost improvements comparable to hardware upgrades, shifting investment priorities for infrastructure teams
  • Agentic AI workloads spanning multiple models, tools, and distributed tasks require orchestration layers that traditional web infrastructure cannot provide, creating demand for specialized serving frameworks
  • Early-mover advantage in inference optimization may compound as software improvements stack across layers, potentially widening cost gaps between optimized and unoptimized deployments

Monitor whether the reported 5x token cost reduction on DeepSeek V4 holds across other models and use cases, or if gains are model-specific. Track adoption rates of NVIDIA's TensorRT-LLM and Dynamo frameworks among competing inference providers to assess whether NVIDIA's software stack becomes industry standard. Watch for similar optimization claims from other hardware vendors and whether they achieve comparable cost reductions.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

DeepSeek Targets $7.5B Funding Close Before Shanghai IPO
TrendingNews

DeepSeek Targets $7.5B Funding Close Before Shanghai IPO

DeepSeek is targeting completion of a $7.5 billion funding round by end-October as it prepares for a Shanghai Stock Exchange IPO. The Chinese AI company's fundraising push is supported by annualized revenue that has reached $1 billion. The timing suggests DeepSeek is accelerating its path to public markets while maintaining aggressive capital raising.

by Claudia Chong· The Information
China Investigates DeepSeek, Moonshot Over Alleged Data Leaks to Anthropic

China Investigates DeepSeek, Moonshot Over Alleged Data Leaks to Anthropic

China's internet regulator is investigating DeepSeek and Moonshot AI following allegations by Anthropic that both companies routed sensitive user data to Claude models without authorization. Anthropic published a 154-page report on September 10 detailing how seven Chinese companies were using Claude illicitly at scale, including an example where DeepSeek relayed requests from engineers building a police surveillance system to Claude. The investigation marks a significant escalation in scrutiny of data practices among Chinese AI firms and raises questions about the security of proprietary AI systems.

by Jing Yang· The Information
DeepSeek Taps Huawei Chips to Sidestep U.S. Export Controls
TrendingNews

DeepSeek Taps Huawei Chips to Sidestep U.S. Export Controls

DeepSeek CEO Liang Wenfeng told investors the company plans to increase use of domestic chips for AI model training, with Huawei expected to begin delivering training chips as early as Q4 2026. The move reflects a coordinated effort by both Chinese companies to circumvent U.S. export controls on advanced semiconductors. DeepSeek is simultaneously closing a second funding round targeting 50 billion yuan ($7.5 billion) at a 500 billion yuan valuation.

by Qianer Liu· The Information
DeepSeek Orders 160,000 Huawei Chips for China Data Center
TrendingNews

DeepSeek Orders 160,000 Huawei Chips for China Data Center

DeepSeek plans to install at least 160,000 Huawei AI chips at a data center in Inner Mongolia, Northern China, according to Bloomberg reporting. The project supports China's broader effort to reduce dependence on Nvidia silicon amid U.S. chip export restrictions. The move signals accelerating domestic chip adoption for large-scale AI infrastructure in China.

by Qianer Liu· The Information