Introducing agentic video understanding with Gemini
Google introduced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, reducing token use by up to 88% and costs by up to 66% while improving accuracy by up to 7%.
Google has launched agentic video understanding for its Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models, enabling dynamic scanning of video segments to improve accuracy while reducing token consumption by up to 88% and costs by up to 66%. The feature is available immediately via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform by setting the API configuration to 'agentic'.
Unlike traditional static video processing, which ingests video at a fixed frame rate, agentic video understanding allows Gemini to actively search, scan, and inspect specific segments across visual frames, audio, and transcripts. This goal-directed approach enables more precise tasks such as sub-second moment retrieval, anomaly detection, and accurate counting, particularly benefiting long-form content like lectures or multi-hour recordings.
Across standard video analysis benchmarks, Gemini models with agentic video understanding demonstrate significant efficiency gains, including up to 66% cost reduction and 88% lower token usage, alongside a 7% accuracy improvement. These gains are most pronounced in long-form video analysis, where static processing often forces trade-offs between cost and detail retention.
The feature is accessible via standard Gemini API token pricing with no additional fees and will soon roll out to all users in the Gemini app across Flash and Flash-Lite models. Additionally, agentic video understanding will power YouTube’s 'Ask YouTube' feature on the video watch page, enhancing answer quality by grounding responses in video content.