arxivcs.CV2026-07-15
Fine-grained CLIP fine-tuning with self-annotated region alignment
Chenyang Zhao, Wei Lin, Antoni B. Chan, Janet H. Hsiao
Contrastive Language-Image Pre-training (CLIP) has been shown to have limitations in its fine-grained dense feature representation, due to its pre-training focusing on matching the whole image to a text description. Considering the large data and computational burden in pre-train…