arxivcs.CVcs.AI2026-07-16
VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation
Xiaoran Xu, Yupeng Wu, Tianyu Xue, Yifan Xu, Xuanran Dong, Xiaoshan Yang, et al.
Object-goal navigation requires an embodied agent to locate and reach an instance of a specified object category in an indoor environment. Recent training-free approaches leverage vision-language models (VLMs) for open-vocabulary semantic reasoning, but are typically evaluated un…