arxivcs.CV2026-07-14
DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery
Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu, Lianshuai Cao, Lei Wang, et al.
As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerged as a key research priority. However, existing…