arxivcs.CVcs.LG2026-07-01
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, Zijian Li, Yifan Shen, Namrata Deka, Harsh Shrivastava, et al.
Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations. These challenges are severely exacerbated by two fundamental properties in the video…