Google has announced an enhancement to the video analysis capabilities of its Gemini AI models, introducing a new agent-based approach. This technological breakthrough significantly reduces the token consumption associated with analyzing video content.
Rather than conventional processing methods that analyze videos on a frame-by-frame basis, the newly announced approach utilizes an intelligent analysis process where an AI agent autonomously assesses the importance of information and processes only the necessary sections. Through this optimization, the company successfully reduced token usage for analysis by up to 88%.
Because video data contains a massive amount of information, high computational costs and token consumption have historically been major challenges when processing it with large language models. The new technology is designed to prevent the excessive consumption of computing resources while maintaining information accuracy, as the AI agent dynamically monitors the video stream to extract and analyze only highly relevant information.
By substantially lowering the cost of video analysis using Gemini, the implementation of this technology is expected to pave the way for the analysis of much longer videos and wider adoption in applications requiring real-time performance.