arxivcs.LGcs.CV2026-06-29
Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
Chungpa Lee, Jihoon Kwon, Kyle Min, Jy-yong Sohn
Vision-language models map images and text into a joint embedding space. However, these embeddings often entangle multiple semantic features, which limits their interpretability and controllability. While sparse autoencoders have emerged as a useful tool for decomposing these emb…