arxivcs.LGcs.CV2026-07-12
On the modality gap and the contrastive loss in multi-modal representation learning
Fabian Mager, Hiba Nassar, Lars Kai Hansen
We study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE formulation with independent encoders. We conduct a u…