arxivcs.CV2026-07-16
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
Xiao Lin, Xiaohu Huang, Kai Han
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs. In this paper, we first r…