arxivcs.CV2026-07-15
CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification
Recent vision models such as CLIP and SAM enable training-free segmentation and semantic encoding for fine-grained classification. A common approach is to compare the representations of segmented image regions with the text prompt embeddings of the corresponding labels. However,…