arxivcs.MMcs.AIcs.LGcs.SD2026-07-01
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
Benjamin Robson, Santeri Mentu, Wenshuai Zhao, Arno Solin
We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg…