VFF - The signal in the noise
News

Weka Extends GPU Memory With Flash Storage to Cut AI Costs

Read original
Share
Weka Extends GPU Memory With Flash Storage to Cut AI Costs

Weka launched NeuralMesh 6, a storage platform designed to reduce GPU memory pressure by caching pre-calculated tokens in cheaper flash storage. The software works alongside Weka's new Wekapod 3 hardware to extend GPU memory using NAND flash at a fraction of the cost. The approach targets enterprises running AI at scale where GPU utilization has become a bottleneck.

  • Weka's NeuralMesh 6 uses flash storage to cache AI model tokens, reducing repeated GPU computation for long context windows and multi-turn conversations
  • The platform supports up to 50,000 tenants per cluster with composable and virtual multi-tenancy options, with provisioning under 30 minutes
  • Unified file and object storage eliminates the need for separate paths and translation layers, claiming roughly two orders of magnitude higher performance than conventional S3
  • Metadata-first replication allows new GPU allocations to become operational within an hour instead of days or weeks for full data migration

GPU memory is the most expensive and fastest-depleting resource in production AI infrastructure. Long context windows force models to repeatedly recompute information, wasting expensive compute cycles. Weka's approach extends GPU memory with cheaper flash storage, potentially allowing organizations to serve more users and deploy workloads faster without purchasing additional GPUs.

For enterprises already operating AI at scale, this reduces inference costs and improves GPU utilization without capital expenditure on new hardware. The faster deployment capability, potentially cutting weeks-long GPU allocation waits to under an hour, directly addresses a competitive constraint that Zvibel says has cost Weka deals.

  • Organizations may be able to defer or reduce GPU capital spending by extending existing memory with cheaper storage alternatives
  • The competitive landscape for AI infrastructure is intensifying, with Dell, NetApp, Pure Storage, and VAST also repositioning toward this category
  • Smaller AI deployments may see limited benefit, making this primarily relevant for enterprises with GPU utilization already constrained

Monitor adoption rates among non-AWS GPU cloud providers that Weka is targeting, particularly Lambda, Nebius, G42, and CoreWeave. Track whether competitors match the claimed performance improvements and multi-tenancy scale. Watch for real-world case studies showing actual cost savings and deployment time reductions compared to traditional GPU scaling.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

DeepSeek Orders 160,000 Huawei Chips for China Data Center
TrendingNews

DeepSeek Orders 160,000 Huawei Chips for China Data Center

DeepSeek plans to install at least 160,000 Huawei AI chips at a data center in Inner Mongolia, Northern China, according to Bloomberg reporting. The project supports China's broader effort to reduce dependence on Nvidia silicon amid U.S. chip export restrictions. The move signals accelerating domestic chip adoption for large-scale AI infrastructure in China.

by Qianer Liu· The Information
Oura Files for IPO on Back of $262M Free Cash Flow

Oura Files for IPO on Back of $262M Free Cash Flow

Oura, the health-tracking ring company, has filed for an IPO. The filing revealed the company generated $262 million in free cash flow over the nine months ending June 30, demonstrating both profitability and scale in the wearable health tech market.

by Martin Peers· The Information
Crusoe Raises $3B at $30B Valuation on Jane Street Deal
TrendingNews

Crusoe Raises $3B at $30B Valuation on Jane Street Deal

Crusoe, a data center developer, has raised $3 billion at a $30 billion valuation, according to reports. The funding round followed the company's reported $13 billion contract with Jane Street. The valuation reflects investor confidence in Crusoe's infrastructure business amid growing demand for data center capacity.

by Marina Temkin· TechCrunch AI
Nvidia to Invest $2.5B in Murati's AI Startup
TrendingNews

Nvidia to Invest $2.5B in Murati's AI Startup

Thinking Machines Lab, the AI startup led by former OpenAI CTO Mira Murati, is raising $5 billion to $6 billion at a valuation of at least $40 billion, with Nvidia expected to contribute roughly $2.5 billion. The funding round reflects Nvidia's strategy to deepen relationships with AI developers and expand its influence across the sector. Andreessen Horowitz, an existing investor with a board seat, is leading investor discussions.

by Stephanie Palazzolo· The Information