arxivcs.MAcs.AIcs.CLcs.CV2026-06-30
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
Qingyun Liu, Jiwen Zhang, Jingyi Hu, Siyuan Wang, Zhongyu Wei
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBench, a multimodal embodied cooperation benchmark with an eva…