arxivcs.CV2026-07-07
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
He Liang, Chenyang Ma, Yiming Zhang, Sangyun Shin, Andrew Markham, Niki Trigoni, et al.
Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the ability to reason over real-world household environments containing multiple interconnected rooms and diverse object categories. We in…