Snowflake's Cortex AI Gateway deploys dynamic routing to cut enterprise AI costs

Snowflake is adding governance-focused controls in Cortex AI Gateway that let enterprises track AI usage, allocate costs, set quotas, and cap spending across teams and agents.
Internal benchmarks tout token-efficiency gains from dynamic routing, including up to 3x token efficiency in some workloads; examples cited include a dbt pipeline achieving up to 3x efficiency and a 25% improvement for the same number of pull requests.
Policy-based routing with a feedback loop: customers can define which models are approved and the tradeoffs to optimize (cost, performance, latency) per workload, while the system continually reevaluates model quality after task completion to adjust routing decisions.
Cortex AI Gateway and dynamic routing sit on a governance layer launched in July 2026; dynamic routing builds on that foundation, whereas previously model selection was handled by a static per-task list.
The move reflects a broader industry shift toward model routing, with Databricks, AWS, Google Cloud, and Nvidia developing their own routing technologies to optimize AI usage and governance.
Snowflake is rolling out Dynamic Model Routing for its Cortex AI Gateway, automatically sending tasks to cheaper or more powerful AI models based on complexity. The feature, now in private preview, aims to cut enterprise AI costs by routing simple tasks to cheaper models and saving frontier models for harder work, according to MarketScreener.
CEO Sridhar Ramaswamy said the next phase of enterprise AI is about economics, not just using the largest model available. Internal tests show dynamic routing can deliver up to 3x token efficiency on some workloads, per Yahoo Finance.
Snowflake uses two routing methods. The first is an advisor pattern — a smaller, cheaper model tries a task first. If it struggles, the system escalates to a larger model. The second uses a history-based classifier that matches new tasks to models that handled similar work well before, according to MarketScreener.
Customers can define which models are approved and set the tradeoffs they want to optimize — cost, performance, or latency. The system then reevaluates model quality after each task and adjusts routing decisions over time. Previously, model selection was handled by a static, fixed list per task.
Snowflake's internal tests show real gains. One example: a dbt pipeline achieved up to 3x token efficiency with dynamic routing. Another workload delivered a 25% improvement for the same number of pull requests. Token efficiency means getting more useful output per dollar spent on AI compute, per MarketScreener.
Ramaswamy told Bloomberg that dynamic routing can cut costs by severalfold on certain workloads. He framed this as a shift in how enterprises think about AI — less about raw model power, more about spending wisely, according to Yahoo Finance.
Dynamic routing builds on a governance layer Snowflake launched in July 2026. That layer lets enterprises track AI usage, allocate costs to specific teams, set quotas, and cap spending across agents and workflows. This is designed for large companies that need to control AI budgets across many departments, per MarketScreener.
Snowflake says governance, context, and model choice are core to enterprise-grade AI — not just speed or price. The company plans to integrate dynamic routing across its flagship products, including Snowflake CoCo and Snowflake CoWork, once the private preview expands.
Snowflake is not alone in this push. Databricks, AWS, Google Cloud, and Nvidia are all building their own model routing tools. The goal across the industry is the same: let AI systems pick the best model automatically, without sacrificing quality, data security, or compliance, according to Yahoo Finance.
Snowflake is also expanding its open-model catalog. DeepSeek V4 Flash 0731 and GLM-5.3 are set to be added to Cortex AI Gateway. These additions give enterprise customers more routing options, especially for high-volume, lower-complexity tasks where cost savings matter most, per MarketScreener.
Publishers
11
Articles
25
Reach
36