CORTEXA
← Browse

Jiwen Zhang

2 papers indexed

arxivcs.ROcs.CV2026-07-04

From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation

Xiangyu Shi, Ruoxi Yang, Wei Tao, Jiwen Zhang, Yanyuan Qiao, Qi Wu

Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation protocols mainly reward stopping within a 3-meter radius and largely ignore the agent's final orienta…

View free PDFSource page
arxivcs.MAcs.AIcs.CLcs.CV2026-06-30

MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

Qingyun Liu, Jiwen Zhang, Jingyi Hu, Siyuan Wang, Zhongyu Wei

Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded environments remains underexplored. To address this gap, we introduce MECoBench, a multimodal embodied cooperation benchmark with an eva…

View free PDFSource page