arxivcs.CV2026-07-06
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
Wei Ao, Lan Wang, Vishnu Naresh Boddeti
The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query confuse their language components. Multimodal retrieval has emerged as a critical component of video understanding, addressing…