arxivcs.CVcs.AI2026-06-28
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo, Shuojin Yang
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks. However, existing benchmarks primarily evaluate whether models c…