arxivcs.CV2026-07-17
IoUPD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
Xiuyuan Zhu, Ke Lu, Hao Wu, Zijin Du, Dongming Zhang, Jian Xue
Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given an image and a referring-expression prompt. While this interface is simple and compatible with instr…