CORTEXA
← Browse
arxivcs.CV2026-07-23

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

Kyeongmo Chae, Jihoon Lee, Sangtae Ahn

This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional deterministic point matching. By modeling both text and video embeddings as Gaussian distributions defined by mean and variance, DAB explicitly accounts for modality-specific uncertainty. We employ a deterministic, diffusion-inspired bridge to iteratively refine text distributions toward their target video distributions through a truncated refinement process. This approach unifies probabilistic embedding and distributional transformation into a cohesive, end-to-end trainable system. To optimize cross-modal similarity, we introduce a distribution-aware contrastive loss based on Kullback-Leibler divergence. Extensive evaluations on MSR-VTT, MSVD, and VATEX benchmarks confirm that DAB significantly outperforms existing probabilistic and diffusion-based baselines, while providing calibrated uncertainty-aware ranking through bridge-induced distributional margins.

View free PDFSource page

Related papers

arxivcs.CVcs.LG2026-06-29

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation

Shihao Zhang, Yunzhi Li, Yuguang Yan, Junzhe Zhang, Wei Zhao, Bohan Wang, et al.

Recent text-to-video (T2V) diffusion models rely heavily on auxiliary reward signals (e.g., via reward models or DPO) to align generated content with human aesthetics and improve realism. These signals, however, incur substantial computational overhead, require costly human annot…

View free PDFSource page
arxivcs.CVcs.LG2026-07-01

MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment

Peiyuan Zhu, Shaoan Xie, Zijian Li, Yifan Shen, Namrata Deka, Harsh Shrivastava, et al.

Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled representations. These challenges are severely exacerbated by two fundamental properties in the video…

View free PDFSource page
arxivcs.CVcs.AIq-bio.NCq-bio.QM2026-07-10

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan, Sepideh Mamooler, Gencer Sumbul, Blair Costelloe, et al.

Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large video-language models (VLMs) have potential to retrieve events of interest by describing them with simple text queries. However, current metho…

View free PDFSource page
arxivcs.CV2026-07-18

When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

Siwei Meng, Yawei Luo, Shu Zhang, Ping Liu

Text-to-video (T2V) generation models have achieved strong visual realism, but improving physical plausibility can come at the cost of semantic consistency with the input text. This tension arises because physical preference is typically determined by comparing dynamics between t…

View free PDFSource page
arxivcs.CV2026-07-21

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

Amber Yijia Zheng, Lu Liu, Raymond A. Yeh, Xi Yin

Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-world data curation is complex and non-trivial, involving clip selection from raw v…

View free PDFSource page