VFF - The signal in the noise
NewsTrending

Apple's Flash-Based Model Architecture Breaks On-Device Memory Ceiling

Read original
Share
Apple's Flash-Based Model Architecture Breaks On-Device Memory Ceiling

Apple announced AFM 3, a new foundation model family developed with Google that includes a 20-billion-parameter on-device model storing weights in NAND flash rather than DRAM. The architecture routes expert selection once per prompt instead of per token, allowing larger models to run locally while staying within consumer device memory constraints. This addresses a fundamental limitation that has kept on-device AI models significantly smaller than cloud alternatives.

  • AFM 3 Core Advanced stores 20B parameters in NAND flash, not DRAM, bypassing the memory ceiling that has limited on-device models
  • Expert routing happens once per prompt, not per token, because NAND-to-DRAM bandwidth cannot support continuous weight swapping
  • Active parameter count scales from 1B to 4B based on task complexity, drawn from the full 20B pool in flash storage
  • Apple developed the architecture with Google and runs server-side models on Nvidia GPUs in Google Cloud within Apple's Private Cloud Compute boundary

On-device AI has been constrained by DRAM capacity, forcing developers to choose between capable cloud models and limited local ones. Apple's flash-based weight storage and per-prompt routing break this constraint, enabling substantially larger models to run locally. This shifts the practical frontier of what on-device AI agents can accomplish without cloud dependency.

Enterprise architects evaluating agentic workloads now have a third option beyond cloud-dependent or limited on-device models. Larger local models reduce latency, improve privacy, and lower cloud compute costs, but deployment viability depends on undisclosed metrics like energy consumption, thermal behavior, and transparent offloading policies that Apple has not yet published.

  • On-device model capacity can now scale to 20B parameters, closing the gap with server-side deployments and enabling more complex local reasoning
  • The per-prompt routing model trades token-level flexibility for memory efficiency, potentially affecting performance on tasks requiring dynamic expert selection across a sequence
  • Apple's undisclosed offloading behavior and lack of energy or thermal profiling data create uncertainty for enterprises planning production deployments
  • The architecture depends on NAND flash speed and DRAM bandwidth characteristics specific to Apple silicon, limiting portability to other platforms

Monitor whether Apple publishes energy, thermal, and bandwidth profiling data needed for production deployment decisions. Watch for third-party benchmarks on real-world agentic workloads and whether transparent offloading to cloud becomes visible to developers and users. Track adoption patterns among enterprise customers evaluating on-device versus hybrid inference strategies.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Saudi Arabia Launches Arabic AI Model With Chinese Partner
TrendingNews

Saudi Arabia Launches Arabic AI Model With Chinese Partner

Humain, Saudi Arabia's state-owned AI company, announced the humain-m3 model, an Arabic language model built on Chinese firm MiniMax's open-source M3 foundation. The model was pre-trained on more than 1 trillion tokens of Arabic content. The development represents a collaboration between Saudi and Chinese AI capabilities focused on Arabic language processing.

by Juro Osawa· The Information
OpenAI's Astra model alarms safety experts with new reasoning technique
News

OpenAI's Astra model alarms safety experts with new reasoning technique

OpenAI's new Astra model employs a technique called 'recurrent depth' that enables reasoning outside the sequential thinking pattern used by most current reasoning models. AI safety experts have raised concerns about this approach. The technique represents a departure from established reasoning architectures in large language models.

by Russell Brandom· TechCrunch AI
Anthropic cuts agent costs 75%, adds enterprise safeguards
TrendingModel Release

Anthropic cuts agent costs 75%, adds enterprise safeguards

Anthropic released Claude Fable 5.1 and Mythos 5.1, its latest large language models, alongside a 75% cost reduction for cached context reads and a new Enterprise Frontier Safeguards security architecture. The release targets enterprise deployment of persistent agents capable of multi-hour problem-solving tasks. Fable 5.1 shows significant benchmark improvements across scientific research, coding, and business workflow tasks, though results are vendor-reported rather than independently verified.

by carl.franzen@venturebeat.com (Carl Franzen)· VentureBeat AI
Chinese AI Model Undercuts US Rivals by 7x on Cost
News

Chinese AI Model Undercuts US Rivals by 7x on Cost

Zhipu's GLM-5.3-Flash model launched on OpenRouter at 7.5 to 25 cents per million tokens (promotional pricing), delivered entirely on Chinese infrastructure. The model scores 57 on Artificial Analysis' intelligence index at roughly nine cents per task, compared to GPT-5.6 Sol at 59 cents and Grok 4.6 at 94 cents, creating significant cost pressure on enterprise AI budgets already strained by unexpected consumption.

· VentureBeat AI