arxivcs.CV2026-07-17
Region-Grounded Vision-Language Learning for Detection-Guided Mammographic Lesion Classification
Zhengbo Zhou, Jiren Li, Dooman Arefan, Margarita Zuley, Shandong Wu
Vision-language models trained with contrastive objectives have shown promise in medical image analysis. However, conventional global image-text alignment is ill-suited for mammography, where diagnostically relevant lesions are spatially localized and occupy only a small fraction…