VFF - The signal in the noise
News

Snowflake adds auto-routing to cut AI query costs up to 3x

Read original
Share
Snowflake adds auto-routing to cut AI query costs up to 3x

Snowflake has launched dynamic model routing in its Cortex AI Gateway, automatically selecting the most cost-effective model for each query rather than using a single fixed model. The company claims the capability can reduce token costs by up to 3x on some workloads by routing simple questions to cheaper models instead of expensive, high-capability ones. The move reflects a broader industry trend toward automated model routing, with competitors including Databricks, AWS, Google Cloud, and Nvidia announcing similar technologies.

  • Snowflake's Cortex AI Gateway now offers auto-routing that selects models based on task complexity and cost, with potential savings up to 3x on token usage
  • Two mechanisms power the routing: an advisor pattern where smaller models attempt tasks first before escalating, and a classifier trained on query history to route straightforward questions to simpler models
  • Routing respects existing access controls and governance boundaries, with customers able to restrict routing to approved model sets or pin specific models
  • Snowflake runs all inference, including open models like DeepSeek-V4-Flash and GLM-5.3, within its security boundary rather than routing to external providers

Enterprise AI deployments face a fundamental tradeoff: using a single capable model for all tasks wastes money on simple queries, while using cheaper models risks quality on complex ones. Automated routing addresses this by matching model capability to task difficulty, reducing unnecessary spending. This capability becomes increasingly important as enterprises scale AI agent deployments and face pressure to optimize cloud spending.

For enterprises running AI agents at scale, model routing directly impacts operational costs and performance. Snowflake's approach ties routing to existing governance and access controls, meaning IT teams can implement cost optimization without rebuilding security infrastructure. The optional nature of auto-routing allows gradual adoption while maintaining control over which models handle sensitive workloads.

  • Model routing is becoming table stakes for AI infrastructure providers, with at least five major vendors now offering some form of automated routing capability
  • Enterprises may need to reconsider their AI spending models, as routing to cheaper models for simple tasks could significantly reduce token costs without sacrificing quality
  • Governance and access control are becoming competitive differentiators, as providers must ensure routing decisions respect data residency, compliance, and permission boundaries

Monitor whether the claimed 3x cost savings hold across diverse enterprise workloads and use cases, as internal testing may not reflect production complexity. Watch for how competitors refine their routing capabilities and whether enterprises adopt auto-routing or prefer manual model selection for control. Track whether open models like DeepSeek and GLM-5.3 become viable for more workloads through better context and routing, potentially shifting spending away from proprietary models.

OneUpAI
OneUp Your Business. Get More Done. OneUp Your Business. Get More Done. OneUp Your Business. Get More Done.
Learn More
Share

Subscribe to the newsletter

The latest stories and analysis, delivered to your inbox.

Free. No spam. Unsubscribe any time.

Related stories

Microsoft Positions Copilot as 'OS for Work'

Microsoft Positions Copilot as 'OS for Work'

Microsoft CEO Satya Nadella presented a strategic rethink of Copilot to enterprise customers at an invite-only event, positioning the AI assistant as the 'OS for work.' The company is integrating coding and agent capabilities directly into Copilot while bringing the full power of Office into the platform for the first time. Nadella is framing this shift as comparable to how Microsoft Office transformed work in the 1980s and 1990s.

by Tom Warren· The Verge AI
Shopify Canvas lets merchants build stores by chatting with AI
TrendingNews

Shopify Canvas lets merchants build stores by chatting with AI

Shopify has launched Canvas, a new site builder that allows merchants to create and customize online stores through conversational AI. The tool uses Shopify's AI agent Sidekick to interpret merchant requests and render changes in real time as users chat. This represents a shift toward natural language interfaces for e-commerce store creation, lowering technical barriers for small business owners.

by Sarah Perez· TechCrunch AI
Vercel's AI Agent Revenue Surge Signals Mainstream Adoption

Vercel's AI Agent Revenue Surge Signals Mainstream Adoption

Vercel, a 11-year-old platform for building and hosting websites and AI applications, is experiencing rapid growth driven by AI coding agents. The company now generates $600 million in annualized revenue, up 148% year-over-year, with coding AI agents accounting for about half of new business, up from less than 3% at the start of the year. Vercel competes with Netlify, Cloudflare, Amazon, and others in a crowded market for developer infrastructure and AI model selection.

by Alix Coutures· The Information
Flow Engineering raises $750M for AI agents in hardware design

Flow Engineering raises $750M for AI agents in hardware design

Flow Engineering, an AI startup focused on applying AI agents to hardware design, raised funding at a $750M valuation. The round was backed by Valor Equity Partners, Atreides Management, and Sequoia Capital. Roelof Botha joined as an angel investor and board member.

by Julie Bort· TechCrunch AI