Identify AI model overuse with User Insights
Cloudflare’s User Insights now provides deeper context on AI model usage, helping teams detect overuse and inefficiencies in AI Gateway traffic through task analysis and model-fit signals.
Cloudflare’s User Insights, launched last month, now offers enhanced context to help teams assess AI model usage. The update introduces task analysis, categorizing conversations by work type such as coding, research, writing, summarization, and data analysis. This allows teams to see whether high-capability models are being used for tasks that require less sophisticated processing, revealing potential inefficiencies in their AI workflows.
The new model overkill view highlights instances where a selected model may exceed the needs of a specific task, such as sending simple formatting requests to a high-capability reasoning model. Teams can investigate which users, agents, or applications are driving this usage and compare metrics like cost, latency, and token usage to determine if adjustments are needed. The view does not automatically replace models but provides data to guide informed decisions about routing rules or workflow changes.
User Insights also introduces a Potential Savings view, which identifies requests that could be handled by faster or less expensive models without compromising output quality. Additionally, the Auto Router, now in public beta, uses task and model-fit signals to automatically route requests to the most appropriate model while balancing cost and performance. This reduces the need for manual routing rules for every workload.
The Auto Router leverages the same task and conversation signals that power User Insights, enabling teams to evaluate model fit across latency, token usage, conversation turns, and total cost. By grouping conversations by task type and complexity, teams can identify patterns such as disproportionate usage of high-capability models for simple tasks and take corrective action. The feature is designed to help organizations optimize AI spending and performance without requiring extensive manual configuration.