arxivcs.CVcs.AI2026-07-01
LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives
Lukas Kuhn, Giuseppe Serra, Randall Balestriero, Florian Buettner
Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted non-contrastive methods. At the same time, the role of vision-language encoders has shifted: they are increasingly deployed not as zero-shot c…