arxivcs.CVcs.AI2026-06-28
SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis
Hang Su, Chao Sun, Zhaofan Li, Wei Hu, Juhua Liu, Bo Du
Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recently emerged, the domain remains particularly challenging due to severe speckle noise, acquisition variability, and subtle anatomica…