arxivcs.CVcs.AI2026-07-15
ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
Kai Chen, Ming Dai, Wenxuan Cheng, Wankou Yang
Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural language expression. However, most advanced methods struggle to balance global context modeling with precise boundary localization. Due…