Release: 2026/09/06 13:41 Reading: 0
Original author:智用AI
Original source:https://www.youtube.com/embed/apmP2Nz4U_0
By dismantling Gemini’s on-demand video understanding mechanism based on Agentic and FPS, long video analysis token consumption is reduced by 88%, and the API supports YouTube links and file uploads. Google has just launched the Agentic video understanding function for the Gemini Flash series models, which marks the shift in video AI processing from "full pre-feeding" to "on-demand retrieval loop". Previously, when models processed long videos, they would blindly ingest all data at a fixed frame rate, causing costs to get out of control and details to be easily missed. Gemini can now autonomously navigate video like an agent, reading keyframes, audio, and transcribed text on demand. In long video analysis, this architecture reduced token consumption by 88% and cost by 66%. For video practitioners, this means you no longer have to face high bills or the loss of fragmented information when dealing with hours of footage or long meeting minutes. This feature is currently available through the Gemini API, supporting file uploads and YouTube links. This article will dismantle the API interaction logic of this dynamic loop, and under what circumstances it is not as efficient as traditional static processing, to help you make accurate decisions. #Gemini #视频ANALYSIS#AgenticAI #big model efficiency#GeminiFlash If this kind of underlying disassembly is useful to you, subscribe + open the small bell, and the new update will be delivered to you as soon as possible. ⏱ Chapter 00:00 When processing a 90-minute video, why are you still using 1 FPS to swallow all the tokens? 00:15 The efficiency bottleneck of traditional model fixed 1 FPS static processing 00:26 The new capabilities of Gemini Flash series launched by Google 00:53 The model is retrieved on demand through the think-action-observation cycle 01:09 Token consumption is reduced by up to 88% and the cost is reduced by 66% 01:24 The standard video benchmark accuracy is increased by up to 7% 01:38 Navigation inference is counted in thought tokens and loaded content is counted in tool-use 01:53 Seamlessly enabled in the API by setting the processing field 02:10 Short videos and low-latency scenarios are not applicable and there are multiple rounds of interaction delays 02:30 The community recognizes the inference speed but the pricing of complaints continues to rise 02:45 There is no open source weight and it is only available for managed APIs 03:00 The best solution for hours-long conference lectures and online class recordings 03:14 A must-see pitfall avoidance and selection list before getting on the bus 🔗 Resources · Original text: https://www.marktechpost.com/2026/09/04/google-agentic-video-understanding-gemini-flash-models/ 🔎 Continue selecting from this video: https://zhiyong.dev/y/apmP2Nz4U_0 --- 🔎 Related AI tools, models and applications: https://zhiyong.dev
Crypto Talk Now
2026-09-21 00:38
Tieu Ai VN
2026-09-21 00:38
Trade with Renato Ulianov
2026-09-21 00:38
Professor Py: AI Engineering
2026-09-21 00:38
Tech with Muthu
2026-09-21 00:38
HBO Max
2026-09-20 22:36
凌云短剧社
2026-09-20 22:20
WCW
2026-09-20 22:20
KarmaStrike Drama
2026-09-20 22:00
Select Currency
US Dollar
USD
Chinese Yuan
CNY
Japanese Yen
JPY
South Korean Won
KRW
New Taiwan Dollar
TWD
Canadian Dollar
CAD
Euro
EUR
Pound Sterling
GBP
Danish Krone
DKK
Hong Kong Dollar
HKD
Australian Dollar
AUD
Brazilian Real
BRL
Swiss Franc
CHF
Chilean Peso
CLP
Czech Koruna KČ
CZK
Singapore Dollar
SGD
Indian Rupee
INR
Saudi Riyal
SAR
Vietnamese Dong
VND
Thai Baht
THB
Select Currency
US Dollar
USD-$
Chinese Yuan
CNY-¥
Japanese Yen
JPY-¥
South Korean Won
KRW -₩
New Taiwan Dollar
TWD-NT$
Canadian Dollar
CAD-$
Euro
EUR - €
Pound Sterling
GBP-£
Danish Krone
DKK-KR
Hong Kong Dollar
HKD- $
Australian Dollar
AUD-$
Brazilian Real
BRL -R$
Swiss Franc
CHF -FR
Chilean Peso
CLP-$
Czech Koruna KČ
CZK -KČ
Singapore Dollar
SGD-S$
Indian Rupee
INR -₹
Saudi Riyal
SAR -SAR
Vietnamese Dong
VND-₫
Thai Baht
THB -฿