arxivcs.IRcs.AIcs.CLcs.CV2026-07-06
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Suhyeong Park, Junha Jung, Jungwoo Park, Jaewoo Kang
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object-…