Google adds agentic video understanding to Gemini
Gemini can now inspect long videos more selectively, cutting token use and cost while improving video analysis quality.
Google DeepMind is adding agentic video understanding to Gemini, a new processing mode meant to make long-form video analysis cheaper and more accurate.
Instead of ingesting a video at a fixed frame rate, Gemini can now decide which parts of a video to search, scan and inspect across frames, audio and transcripts. Google says this helps with tasks where one-frame-per-second processing misses important details or burns too many tokens on irrelevant footage.
The company says the mode reduces token consumption by up to 88%, cuts analysis costs by up to 66% and improves quality by up to 7% on standard video understanding benchmarks. Google frames the biggest gains around long videos such as lectures, how-to recordings, surveillance-style clips and multi-hour archives.
For developers, the feature is available on Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It works with uploaded videos and YouTube videos, uses standard Gemini API pricing and can be enabled by setting video processing to agentic.
Google also says the same capability will come to the Gemini app and will later power Ask YouTube, where users ask questions about a video while watching it.
Sources
- Google DeepMindblog.google