arxivcs.CVcs.AI2026-07-03
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning
Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu
Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We aim to learn representations whose matching is stabl…