arxivcs.CV2026-07-10
SigLIP-HD by Fine-to-Coarse Supervision
Lihe Yang, Zhen Zhao, Hengshuang Zhao
High-quality visual representation is a long-standing pursuit in computer vision. In the context of multimodal LLMs (MLLMs), feeding higher-resolution images can produce more fine-grained visual tokens. However, it introduces additional computational and design complexity, due to…