Snowflake adds auto-routing to cut AI query costs up to 3x

Snowflake has launched dynamic model routing in its Cortex AI Gateway, automatically selecting the most cost-effective model for each query rather than using a single fixed model. The company claims the capability can reduce token costs by up to 3x on some workloads by routing simple questions to cheaper models instead of expensive, high-capability ones. The move reflects a broader industry trend toward automated model routing, with competitors including Databricks, AWS, Google Cloud, and Nvidia announcing similar technologies.
TL;DR
- Snowflake's Cortex AI Gateway now offers auto-routing that selects models based on task complexity and cost, with potential savings up to 3x on token usage
- Two mechanisms power the routing: an advisor pattern where smaller models attempt tasks first before escalating, and a classifier trained on query history to route straightforward questions to simpler models
- Routing respects existing access controls and governance boundaries, with customers able to restrict routing to approved model sets or pin specific models
- Snowflake runs all inference, including open models like DeepSeek-V4-Flash and GLM-5.3, within its security boundary rather than routing to external providers
Why It Matters
Enterprise AI deployments face a fundamental tradeoff: using a single capable model for all tasks wastes money on simple queries, while using cheaper models risks quality on complex ones. Automated routing addresses this by matching model capability to task difficulty, reducing unnecessary spending. This capability becomes increasingly important as enterprises scale AI agent deployments and face pressure to optimize cloud spending.
Business Impact
For enterprises running AI agents at scale, model routing directly impacts operational costs and performance. Snowflake's approach ties routing to existing governance and access controls, meaning IT teams can implement cost optimization without rebuilding security infrastructure. The optional nature of auto-routing allows gradual adoption while maintaining control over which models handle sensitive workloads.
Key Implications
- Model routing is becoming table stakes for AI infrastructure providers, with at least five major vendors now offering some form of automated routing capability
- Enterprises may need to reconsider their AI spending models, as routing to cheaper models for simple tasks could significantly reduce token costs without sacrificing quality
- Governance and access control are becoming competitive differentiators, as providers must ensure routing decisions respect data residency, compliance, and permission boundaries
What to Watch
Monitor whether the claimed 3x cost savings hold across diverse enterprise workloads and use cases, as internal testing may not reflect production complexity. Watch for how competitors refine their routing capabilities and whether enterprises adopt auto-routing or prefer manual model selection for control. Track whether open models like DeepSeek and GLM-5.3 become viable for more workloads through better context and routing, potentially shifting spending away from proprietary models.
Subscribe to the newsletter
The latest stories and analysis, delivered to your inbox.
Free. No spam. Unsubscribe any time.


