arxivcs.CV2026-07-15
AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization
Yiyang Yao, Shanglin Liu, Jianming Lv, Chengjun Wang, Jinyi Li, Yuchan Jie, et al.
Contrastive Language-Image Pretraining learns a shared representation space through large-scale contrastive learning. However, existing methods that enforce global consistency regularization overlook a key challenge: the inherent information asymmetry between images and text: cap…