HeadlinesBriefing favicon HeadlinesBriefing.com

Google DeepMind Launches Agentic Video Understanding for Gemini

Google DeepMind Blog •
×

Google DeepMind has launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models, cutting token consumption by up to 88%, reducing costs by up to 66%, and boosting accuracy by up to 7%. Unlike static processing at fixed frame rates, this feature lets Gemini dynamically search, scan, and inspect video segments across visual frames, audio, and transcripts via an agentic loop.

Benchmarks show the efficiency gains are most pronounced on long-form content — from 10-minute guides to multi-hour recordings — where developers previously faced trade-offs between high token costs and dropped details. Gemini 3.7 Flash with agentic understanding now sits at the accuracy-to-cost pareto frontier for video analysis.

Key capabilities include sub-second moment retrieval for precise editing, needle-in-a-haystack search across multi-hour videos, anomaly detection via high-FPS resampling, and accurate counting of actions and objects. Early access partners reported strong results. The feature is available today via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform for video uploads and YouTube videos.