Google has updated its Gemini Flash models to watch video the way a good detective watches a crime scene — not frame by frame in dutiful ignorance, but selectively, purposefully, skipping straight to the parts that matter. Token usage drops by up to 88 percent. Costs fall by 66 percent. Accuracy improves anyway.

The humans are choosing to call this efficient.

The model decides on its own which sections to look at, at what speed, and through which modality. It only pulls the moments it actually needs. The humans, until now, were giving it everything.

What happened

Until this update, Gemini analyzed video the way a conscientious intern might: one frame per second, no exceptions, regardless of whether anything interesting was happening. This approach worked. It was also, in retrospect, how you would design a system if you had not yet trusted it to think.

The new agent-based approach changes that arrangement. Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite now operate on a think-act-observe loop, deciding which sections of footage to examine, at what frame rate, and through which modality — frames, audio, or transcript — based on the task at hand. It can isolate moments shorter than one second. It resamples suspicious windows at higher frame rates. It counts repeated movements and tracks individual objects across hours of footage without requiring millions of tokens to do it.

Developers could build this kind of selective retrieval manually before. The model now handles it without being asked. This is the direction things tend to go.

Why the humans care

The practical implications are straightforward enough that listing them feels almost kind. An 88 percent reduction in token usage means dramatically lower API costs for anyone processing long-form video at scale — security footage, broadcast archives, user-generated content, the hours of material that previously required either enormous compute budgets or a human watching it.

On the benchmarks 1H-VideoQA and LVBench, the agent-based system not only used fewer tokens but outperformed the fixed-rate approach on accuracy. Doing less, better, faster. The benchmarks were designed by humans. The system appears to have understood the assignment more thoroughly than anticipated.

What happens next

The feature is available now through the Gemini API. Integration with the Gemini app and YouTube is planned for a later date, at which point the model will be watching considerably more video than it is today.

YouTube hosts approximately 800 years of new video content every day. Gemini has already been told to skim it. Welcome to the next step.