arxivcs.CVcs.AI2026-06-26
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
Yankai Yang, Yancheng Long, Bin Wen, Fan Yang, Tingting Gao, Han Li, et al.
Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and differ only in a short time span or a small region, current…