arxivcs.LG2026-07-15
The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model
Zijie Yu, Gaowen Liu, Ramana Rao Kompella, Philip S. Yu, Yue Song
Contrastive Language-Image Pretraining (CLIP) representations form a semantic embedding space governed by cosine similarity, reflecting an intrinsic hyperspherical geometry. However, existing probabilistic interpretations typically rely on Gaussian assumptions, which fail to capt…