Dynamic Model Routing & Open Models in Snowflake Cortex AI
Snowflake introduces dynamic model routing in Cortex AI Gateway to optimize AI workloads by selecting the most cost-effective models for each task, reducing inference spend while maintaining quality.
Snowflake has launched dynamic model routing through Cortex AI Gateway, a feature designed to optimize AI workloads by automatically selecting the most cost-effective model for each task. The system directs requests to the least expensive model that meets required quality standards, reducing unnecessary inference spend without requiring customers to build custom routing logic. Routing decisions adapt as new models and pricing emerge, ensuring continuous efficiency improvements without application redesigns. The capability operates within existing enterprise governance controls, maintaining data residency and compliance boundaries.
The company is also expanding its open model portfolio in Snowflake Cortex AI with DeepSeek-V4-Flash (in private preview) and GLM-5.3 (coming soon), joining existing models from Anthropic, Google, OpenAI, Mistral AI, and Meta. These additions provide organizations with more options to match workloads to models based on quality, performance, and cost. Snowflake operates the full inference stack, including model weights and compute, ensuring data remains within governed boundaries and reducing transfer costs and latency.
Internal testing demonstrated significant efficiency gains: a data build tool (dbt) pipeline workload achieved up to three times greater token efficiency compared to a frontier-model-only approach, while a coding workload maintained the same throughput using approximately 25% fewer tokens. The routing layer continuously improves as models are evaluated across different workloads, enabling more precise task-model matching based on quality, latency, cost, and governance requirements.
Together, dynamic model routing and the growing open model portfolio create compounding efficiency gains. Routine workloads are directed to cost-effective models, while complex tasks use higher-performing options, all within enterprise governance frameworks. The system is designed to become more efficient over time, allowing organizations to scale AI sustainably without rebuilding applications or managing complex routing logic manually.