The Specialty News
AI

Google Gives Gemini a Remote Control, and AI Video Costs Crash 88%

By turning models from passive stenographers into active investigators, DeepMind just solved the biggest bottleneck in multimodal computing.

By The Specialty News DeskEdited by 4 min read
Google Gives Gemini a Remote Control, and AI Video Costs Crash 88%
Photo: deepmind.google

Hand a traditional AI model a video, and it acts like a hostage strapped to a chair, forced to watch every single frame at exactly one second per frame. Google DeepMind just cut the ropes. With a new update across its Gemini models, the AI now operates like a human investigator armed with a remote control—scanning the transcript first, fast-forwarding through dead air at 0.1 frames per second, and slowing down to catch fast action. It is a fundamental shift in how machines process reality, turning video analysis from a cost-prohibitive luxury into an accessible default.

The 100,000-Token Tollbooth

Yesterday, feeding a 60-minute meeting recording into a multimodal model meant swallowing 3,600 individual frames and burning through more than 100,000 tokens. It was a brute-force approach. To answer a simple question about what the CEO said in minute 42, the AI had to pay the computational toll for the other 59 minutes of throat-clearing and screen-sharing.

By switching a single API parameter to "agentic," developers now flip the model from passive to active. Instead of ingesting the entire video file at a fixed rate, Gemini runs an internal loop to think, act, and observe. It reads the transcript like a map, scrubs to the precise 10-second slice it actually needs, and only turns on the audio track if acoustic cues matter.

It is the computational equivalent of reading a book's index instead of reading cover-to-cover just to find one name. The model escapes the noise, driving accuracy up by 7% precisely because it isn't drowning in irrelevant data. But this surgical efficiency doesn't just improve test scores—it flips the underlying economics of the AI industry.

Scuttling the Middleware Market

Scuttling the Middleware Market
Photo: searchenginejournal.com

For the past year, processing long-form media was the ultimate bottleneck. It locked smaller builders out of complex video analysis, trapping video-RAG tools and automated sports analytics platforms in the prototype phase. The prohibitive cost of tokens forced builders to rely on a cottage industry of startups that built proprietary skimming middleware to compress video before the AI saw it.

Rather than spending 200k tokens processing a long video, Gemini can now use tools to intelligently process them based on the prompt... leading to 88% fewer tokens.Omar Sanseviero

Google just baked that entire middleware industry directly into the base API. By making agentic reasoning a native feature, architects like DeepMind's Rohan Doshi and Mario Lučić are collapsing the technology stack.

The startups that banked on token costs remaining high are suddenly obsolete, while independent developers building the next generation of transcription agents just received a massive margin boost. Yet, for all its structural elegance, handing an AI the remote control introduces a new kind of risk.

The TikTok Trap

The headline-grabbing efficiency is a best-case scenario drawn from benchmarks where hours of footage are actually irrelevant to the prompt. If you hand Gemini a dense, 30-second TikTok clip where every frame is packed with visual cues, there is no dead air to skip.

In fact, Google quietly advises developers to keep static processing turned on for clips under a few minutes. The computational overhead required to boot up the agentic loop can actually slow down response times for quick queries. Skeptics also point to the lingering fear of the leaky loop, wondering if a model might wrongly decide to fast-forward past a crucial, silent visual clue just because it wasn't flagged in the transcript first.

DeepMind isn't waiting for the industry to settle the debate. In the coming months, this exact agentic engine will power an "Ask YouTube" feature directly on watch pages, bringing surgical video analysis to billions of consumers. We spent the last two years teaching AI how to watch video. Now, we are teaching it what to ignore.

Passive vs Agentic Video AI

A visual summary of this story

The Brief

Stay curious

AI and technology: what changes and why it matters.
Your daily selection, in English or Spanish.

Free forever. Unsubscribe anytime.

Conversation

Start the conversation

No account needed. Comments are checked automatically — keep it civil.

More stories

Keep reading