CORTEXA
← Browse
openalexCybersecurity2026-07-24Cited by 0

Temporal modality reliability and uncertainty-aware alignment for text-video retrieval

B Yang, Yun Cao, H H. Chen, Hong Zhang

Abstract The rapid growth of multimodal content has introduced new challenges in cybersecurity, particularly in scenarios such as misinformation detection, multimedia forensics, and open-source intelligence. In these settings, verifying the consistency between textual descriptions and video content is critical, yet remains challenging due to noisy, incomplete, or even misleading multimodal signals. Text-to-video retrieval provides a important capability for cross-modal alignment. However, existing approaches often overlook a key issue: the usefulness of different modalities varies over time and depends on the query. In particular, audio signals can be informative in certain segments (e.g., speech) but misleading in others (e.g., background noise), making uniform fusion unreliable in security-critical scenarios. In this work, we propose Temporal Uncertainty-aware Retrieval (TUR), a unified framework that models text-video alignment from two complementary perspectives: temporal modality reliability and alignment uncertainty. TUR dynamically estimates the contribution of multimodal signals over time and adapts text representations according to cross-modal agreement, enabling more stable retrieval. Extensive experiments on MSR-VTT, DiDeMo, VATEX, and LSMDC demonstrate that TUR consistently outperforms prior methods. Further analysis shows that TUR achieves improved temporal grounding, more stable similarity estimation, and enhanced interpretability, which are desirable properties for security-sensitive applications.

View free PDFSource page