VFF - The signal in the noise
News

Weka Extends GPU Memory With Flash Storage to Cut AI Costs

Read original
Share
Weka Extends GPU Memory With Flash Storage to Cut AI Costs

Weka launched NeuralMesh 6, a storage platform designed to reduce GPU memory pressure by caching pre-calculated tokens in cheaper flash storage. The software works alongside Weka's new Wekapod 3 hardware to extend GPU memory using NAND flash at a fraction of the cost. The approach targets enterprises running AI at scale where GPU utilization has become a bottleneck.

  • Weka's NeuralMesh 6 uses flash storage to cache AI model tokens, reducing repeated GPU computation for long context windows and multi-turn conversations
  • The platform supports up to 50,000 tenants per cluster with composable and virtual multi-tenancy options, with provisioning under 30 minutes
  • Unified file and object storage eliminates the need for separate paths and translation layers, claiming roughly two orders of magnitude higher performance than conventional S3
  • Metadata-first replication allows new GPU allocations to become operational within an hour instead of days or weeks for full data migration

GPU memory is the most expensive and fastest-depleting resource in production AI infrastructure. Long context windows force models to repeatedly recompute information, wasting expensive compute cycles. Weka's approach extends GPU memory with cheaper flash storage, potentially allowing organizations to serve more users and deploy workloads faster without purchasing additional GPUs.

For enterprises already operating AI at scale, this reduces inference costs and improves GPU utilization without capital expenditure on new hardware. The faster deployment capability, potentially cutting weeks-long GPU allocation waits to under an hour, directly addresses a competitive constraint that Zvibel says has cost Weka deals.

  • Organizations may be able to defer or reduce GPU capital spending by extending existing memory with cheaper storage alternatives
  • The competitive landscape for AI infrastructure is intensifying, with Dell, NetApp, Pure Storage, and VAST also repositioning toward this category
  • Smaller AI deployments may see limited benefit, making this primarily relevant for enterprises with GPU utilization already constrained

Monitor adoption rates among non-AWS GPU cloud providers that Weka is targeting, particularly Lambda, Nebius, G42, and CoreWeave. Track whether competitors match the claimed performance improvements and multi-tenancy scale. Watch for real-world case studies showing actual cost savings and deployment time reductions compared to traditional GPU scaling.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

LTX-2.5 Generates Video Faster Than Real-Time, Pushes Open Weights Forward

LTX-2.5 Generates Video Faster Than Real-Time, Pushes Open Weights Forward

LTX released LTX-2.5, an open-weights video generation model that produces 10-second clips in 6.8 seconds on Nvidia GB200 chips, with native multishot support and improved quality. The model is available free for organizations under $10 million ARR on Hugging Face, ComfyUI, and via API. LTX claims 33 million downloads across its model family and reports a 67% win rate in blind quality tests against competing models.

by carl.franzen@venturebeat.com (Carl Franzen)· VentureBeat AI
OpenAI Robotics Lead Joins Anthropic
TrendingNews

OpenAI Robotics Lead Joins Anthropic

Caitlin Kalinowski, former head of robotics at OpenAI, has joined Anthropic as a member of technical staff focused on research. The hire signals Anthropic's continued investment in robotics capabilities, following the company's release of robotics research last month. Kalinowski's move represents a notable talent shift between two of the leading AI research organizations.

by Rocket Drew· The Information
Microsoft Bets Big on Homegrown AI Chips Despite Slow Start

Microsoft Bets Big on Homegrown AI Chips Despite Slow Start

Microsoft is planning to dramatically scale production of its next-generation Maia 300 AI chip, with plans to order over 300,000 units from Taiwan Semiconductor Manufacturing Co. for 2027 delivery. The move comes despite weak adoption of the current Maia 200 generation and represents a significant bet that cloud customers like Anthropic will adopt the new design. Microsoft plans to publicly unveil the Maia 300 this fall.

by Qianer Liu· The Information
Firebird Opens CIS AI Factory in Armenia, Eyes 2-Gigawatt Expansion

Firebird Opens CIS AI Factory in Armenia, Eyes 2-Gigawatt Expansion

Firebird, an emerging AI cloud provider, opened the CIS region's largest AI factory in Armenia with support from NVIDIA and Dell Technologies. The facility will deploy over 70,000 NVIDIA Rubin and Blackwell GPUs and 300 megawatts of capacity by end of 2027, part of Firebird's broader 2-gigawatt infrastructure roadmap across frontier markets. The opening was attended by Armenia's prime minister, Kazakhstan's deputy prime minister, and the U.S. chargé d'affaires, signaling geopolitical interest in regional AI infrastructure development.

by Rev Lebaredian· NVIDIA Blog (AI)