arxivcs.CV2026-07-01
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
Ji Ha Jang, Hayeon Kim, Chulwon Lee, Junghun James Kim, Se Young Chun
CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions (>77 tokens) due to absolute positional encoding and pretraining on short captions. In long contexts, sentences are often reorde…