arxivcs.CV2026-07-14
CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models
Lin Peng, Cong Wan, Zeyu Guo, SongLin Dong, Yihong Gong
Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally consistent reasoning. We present CoRe, a unified framework for this problem. CoRe includes: (i) CoRe-…