Cut your AI spend with AI Gateway's Auto Router
Cloudflare launches AI Gateway’s Auto Router in public beta to automate model selection and reduce AI spending by up to 30% without user intervention.
Cloudflare introduces Auto Router in public beta through AI Gateway, enabling automatic model selection for AI requests to optimize cost and performance. Users set their model to cloudflare/auto, allowing the system to route each request to the most suitable model without manual intervention. The tool aims to address common inefficiencies in AI adoption, where organizations often overpay for high-end models on tasks that don’t require them. Early internal testing at Cloudflare showed cost savings of up to 30% compared to using only frontier models like OpenAI Sol and Anthropic Claude Opus.
AI Gateway’s Auto Router is designed to act as a control plane for organizations managing internal AI deployments, building on existing cost-control features like budgets, spend limits, and identity-aware analytics. Previously, users manually selected models in tools like OpenCode and Claude Code, often choosing overpowered models for simple tasks. The Auto Router eliminates this need by intelligently routing requests based on task requirements, ensuring users retain access to high-capability models when necessary while reducing unnecessary spending.
The router evaluates requests by analyzing conversation context, task complexity, ambiguity, stakes, and dependence on prior context using a multi-head classification model running on Workers AI. It assigns probabilities across 14 task categories and rates requests on four dimensions to determine the best-fit model. The system combines these signals with benchmark results and token pricing to estimate utility, prioritizing cost efficiency for straightforward tasks while allowing stronger models to handle more demanding workloads.
Auto Router accounts for additional cost factors beyond token prices, such as cache read/write costs during model switching in long agentic sessions. It applies a switching penalty that grows with context length, ensuring the system only switches models when the long-term savings justify the cost. The tool also considers the loss of reasoning tokens during switches, which may require the new model to regenerate output at higher prices. Future updates aim to further refine these calculations to minimize total cost while maintaining performance.