Weka Extends GPU Memory With Flash Storage to Cut AI Costs

Weka launched NeuralMesh 6, a storage platform designed to reduce GPU memory pressure by caching pre-calculated tokens in cheaper flash storage. The software works alongside Weka's new Wekapod 3 hardware to extend GPU memory using NAND flash at a fraction of the cost. The approach targets enterprises running AI at scale where GPU utilization has become a bottleneck.
TL;DR
- Weka's NeuralMesh 6 uses flash storage to cache AI model tokens, reducing repeated GPU computation for long context windows and multi-turn conversations
- The platform supports up to 50,000 tenants per cluster with composable and virtual multi-tenancy options, with provisioning under 30 minutes
- Unified file and object storage eliminates the need for separate paths and translation layers, claiming roughly two orders of magnitude higher performance than conventional S3
- Metadata-first replication allows new GPU allocations to become operational within an hour instead of days or weeks for full data migration
Why It Matters
GPU memory is the most expensive and fastest-depleting resource in production AI infrastructure. Long context windows force models to repeatedly recompute information, wasting expensive compute cycles. Weka's approach extends GPU memory with cheaper flash storage, potentially allowing organizations to serve more users and deploy workloads faster without purchasing additional GPUs.
Business Impact
For enterprises already operating AI at scale, this reduces inference costs and improves GPU utilization without capital expenditure on new hardware. The faster deployment capability, potentially cutting weeks-long GPU allocation waits to under an hour, directly addresses a competitive constraint that Zvibel says has cost Weka deals.
Key Implications
- Organizations may be able to defer or reduce GPU capital spending by extending existing memory with cheaper storage alternatives
- The competitive landscape for AI infrastructure is intensifying, with Dell, NetApp, Pure Storage, and VAST also repositioning toward this category
- Smaller AI deployments may see limited benefit, making this primarily relevant for enterprises with GPU utilization already constrained
What to Watch
Monitor adoption rates among non-AWS GPU cloud providers that Weka is targeting, particularly Lambda, Nebius, G42, and CoreWeave. Track whether competitors match the claimed performance improvements and multi-tenancy scale. Watch for real-world case studies showing actual cost savings and deployment time reductions compared to traditional GPU scaling.
Subscribe to the newsletter
The latest stories and analysis, delivered to your inbox.
Free. No spam. Unsubscribe any time.