arxivcs.RO2026-07-01
Where Am I? Semantic Map Grounding via Vision-Language Models for Multi-Modal Localization
Suraj Borate, Aarav Shah, Madhu Vadali
We address robot localization in GPS-denied indoor environments by reframing it as a semantic reasoning task rather than a geometric estimation problem. Motivated by how humans localize using object-level cues and labeled maps, we ask whether a vision-language model, given a fron…