arxivcs.CVcs.AI2026-07-07
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Daniel Shalam, Emanuel Ben Baruch, Avi Ben Cohen, Tal Remez
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events, but they hallucinate these regions prolifically. The model's own token log-probabilities are nearly uninformative: they conflate grounding…