VFF - The signal in the noise
News

NVIDIA Opens Storage APIs as AI Demands Overwhelm Infrastructure

Read original
Share
NVIDIA Opens Storage APIs as AI Demands Overwhelm Infrastructure

NVIDIA is addressing a critical bottleneck in AI infrastructure by open sourcing its cuFile APIs and advancing storage optimization initiatives. As AI workloads demand massive datasets and context windows that exceed system memory, storage systems must handle thousands of concurrent GPU-initiated requests while performing encryption, compression, and data verification. NVIDIA's Vera CPU delivers up to 3.21x higher throughput than x86 alternatives in compression and encryption pipelines, while the open sourcing of cuFile enables GPUs to access storage directly in microseconds rather than minutes.

  • NVIDIA open sourced cuFile APIs, enabling GPUs to read and write directly to storage with microsecond latency
  • NVIDIA Vera CPU shows 3.21x higher throughput than x86 CPUs in two-stage compression and encryption pipelines
  • Storage is shifting from passive data repository to active component in the data path for AI workloads
  • NVIDIA launched Storage-Next initiative with storage makers, controller vendors, and standards bodies to align GPU-driven storage behavior

AI systems now generate thousands of concurrent storage requests that traditional architectures cannot efficiently handle. Storage operations like encryption, compression, and verification have become critical bottlenecks. Solving this requires rethinking the memory versus storage tradeoff, which has shifted from minute-scale access times to microsecond-scale performance on modern GPUs.

Organizations deploying AI agents at scale face infrastructure costs and performance constraints driven by storage inefficiency. Open sourcing cuFile and advancing storage optimization reduces the compute overhead needed to manage data access, lowering total cost of ownership while improving throughput. Companies building AI infrastructure now have standardized, interoperable tools to address storage as a critical performance lever.

  • Storage infrastructure is no longer a cost-optimization decision but a performance-critical component of AI systems
  • Open sourcing cuFile with Google, Intel, Meta, and NVIDIA as maintainers signals industry-wide standardization around GPU-native storage access
  • The 3.21x throughput improvement of Vera CPUs suggests specialized hardware for storage operations will become standard in AI deployments

Monitor adoption of cuFile APIs across storage vendors and cloud providers to gauge industry standardization. Track performance benchmarks from Storage-Next initiative participants to see if specialized storage controllers become mainstream. Watch for announcements from major cloud providers integrating these storage advancements into their AI infrastructure offerings.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Runware launches portable data center pod for AI inference
TrendingNews

Runware launches portable data center pod for AI inference

AI infrastructure company Runware announced the launch of Sonic Inference Pod, a modular data center designed for portable deployment. The product represents an effort to address the growing demand for distributed AI computing infrastructure. The announcement signals a shift toward more flexible, deployable data center solutions in the AI hardware space.

by Dominic-Madori Davis· TechCrunch AI
CyrusOne Prepares Major Data Center IPO

CyrusOne Prepares Major Data Center IPO

CyrusOne, a data center operator owned by KKR and BlackRock's Global Infrastructure Partners, is preparing for what could be one of the largest IPOs next year by interviewing investment banks. The company was taken private in early 2022 for $15 billion. The IPO would allow the private equity owners to cash out while enabling CyrusOne to pay down debt accumulated during data center expansion.

by Valida Pau· The Information
Xsight Raises $300M as Investors Back GPU Networking Infrastructure

Xsight Raises $300M as Investors Back GPU Networking Infrastructure

Xsight, an Israeli server networking and storage chip startup, raised $300 million at a $2.8 billion valuation in its first major funding round in five years. Led by Fidelity Investments with participation from Atreides Management, Valor Equity Partners, Battery Ventures, and Intel Capital, the round reflects investor appetite for networking infrastructure that connects GPUs in data centers. The funding follows recent rounds for competing networking startups Eliyan and Upscale AI, signaling a broader market opportunity beyond GPU chips themselves.

by Phoebe Liu· The Information
Google DeepMind Releases Gemini Robotics 2 for Whole-Body Robot Control
TrendingModel Release

Google DeepMind Releases Gemini Robotics 2 for Whole-Body Robot Control

Google DeepMind introduced Gemini Robotics 2, a suite of AI models designed to give robots whole-body control, dexterous manipulation, and multi-robot collaboration capabilities. The system includes three models: a vision-language-action model for motor control, an embodied reasoning model for planning and communication, and an on-device model optimized for fast adaptation to new robot bodies. Early-access partners can now deploy these models on humanoid and bi-arm robots to perform complex, multi-step tasks in unstructured environments.

· Google Deepmind