VFF - The signal in the noise
News

Weka Extends GPU Memory With Flash Storage to Cut AI Costs

Read original
Share
Weka Extends GPU Memory With Flash Storage to Cut AI Costs

Weka launched NeuralMesh 6, a storage platform designed to reduce GPU memory pressure by caching pre-calculated tokens in cheaper flash storage. The software works alongside Weka's new Wekapod 3 hardware to extend GPU memory using NAND flash at a fraction of the cost. The approach targets enterprises running AI at scale where GPU utilization has become a bottleneck.

  • Weka's NeuralMesh 6 uses flash storage to cache AI model tokens, reducing repeated GPU computation for long context windows and multi-turn conversations
  • The platform supports up to 50,000 tenants per cluster with composable and virtual multi-tenancy options, with provisioning under 30 minutes
  • Unified file and object storage eliminates the need for separate paths and translation layers, claiming roughly two orders of magnitude higher performance than conventional S3
  • Metadata-first replication allows new GPU allocations to become operational within an hour instead of days or weeks for full data migration

GPU memory is the most expensive and fastest-depleting resource in production AI infrastructure. Long context windows force models to repeatedly recompute information, wasting expensive compute cycles. Weka's approach extends GPU memory with cheaper flash storage, potentially allowing organizations to serve more users and deploy workloads faster without purchasing additional GPUs.

For enterprises already operating AI at scale, this reduces inference costs and improves GPU utilization without capital expenditure on new hardware. The faster deployment capability, potentially cutting weeks-long GPU allocation waits to under an hour, directly addresses a competitive constraint that Zvibel says has cost Weka deals.

  • Organizations may be able to defer or reduce GPU capital spending by extending existing memory with cheaper storage alternatives
  • The competitive landscape for AI infrastructure is intensifying, with Dell, NetApp, Pure Storage, and VAST also repositioning toward this category
  • Smaller AI deployments may see limited benefit, making this primarily relevant for enterprises with GPU utilization already constrained

Monitor adoption rates among non-AWS GPU cloud providers that Weka is targeting, particularly Lambda, Nebius, G42, and CoreWeave. Track whether competitors match the claimed performance improvements and multi-tenancy scale. Watch for real-world case studies showing actual cost savings and deployment time reductions compared to traditional GPU scaling.

Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

AMD commits $5B to Anthropic, will supply 2GW of AI chips
TrendingNews

AMD commits $5B to Anthropic, will supply 2GW of AI chips

AMD announced a commitment of up to $5 billion in investment to Anthropic and will supply the AI company with up to 2 gigawatts of its Instinct MI450 AI GPUs using the Helios rack-scale system. The first gigawatt is scheduled for deployment in the first half of 2027. This deal expands Anthropic's infrastructure partnerships, which already include agreements with SpaceX, TeraWulf, Google, Broadcom, and Amazon.

by Emma Roth· The Verge AI
NVIDIA Open-Sources Medical Robotics Simulation Framework

NVIDIA Open-Sources Medical Robotics Simulation Framework

NVIDIA has open-sourced the Medical Physics Simulation framework, a GPU-accelerated tool within NVIDIA Isaac for Healthcare that enables medical robotics developers to simulate anatomy-device interactions, generate training scenarios, and test robot behavior in virtual environments before physical testing. The framework combines classical physics simulation with generative AI to model complex surgical scenarios like vascular procedures with catheters and guidewires. By running hundreds of parallel simulations on GPU hardware, developers can reduce training time from over five hours to under two minutes, addressing a major bottleneck in healthcare robotics development.

by David Niewolny· NVIDIA Blog (AI)
Wistron Opens Fort Worth AI Superchip Plant, Part of $500B U.S. Push

Wistron Opens Fort Worth AI Superchip Plant, Part of $500B U.S. Push

Wistron opened its first U.S. manufacturing facility in Fort Worth, Texas, a 324,000-square-foot plant producing NVIDIA superchips for AI systems. The $700 million facility currently operates two manufacturing cells producing the GB300 Grace Blackwell Ultra Superchip and will produce the Vera Rubin Superchip, with plans to scale to tens of thousands of boards per month. The plant has created over 500 jobs with expansion to 1,000 planned by year-end, representing part of NVIDIA's broader $500 billion commitment to U.S. advanced AI manufacturing.

by NVIDIA Writers· NVIDIA Blog (AI)
NVIDIA Spectrum-6 Targets AI's New Bottleneck: Network Performance

NVIDIA Spectrum-6 Targets AI's New Bottleneck: Network Performance

NVIDIA has released Spectrum-6, a 102.4-terabit-per-second Ethernet switch system that doubles the capacity of previous-generation systems and is designed for gigascale AI infrastructure. The switch is part of the NVIDIA Vera Rubin platform and is being deployed by CoreWeave, Microsoft, Nebius, SpaceX AI, and Tesla to improve coordination across hundreds of thousands of GPUs and CPUs in AI factories. Spectrum-6 addresses a fundamental constraint in large-scale AI training and inference, where network performance rather than individual GPU performance becomes the limiting factor.

by Scot Schultz· NVIDIA Blog (AI)