CORTEXA
← Browse

Danfei Xu

5 papers indexed

arxivcs.CVcs.RO2026-07-09

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

Utkarsh A. Mishra, Yongxin Chen, Danfei Xu, Yang Liu, Xi Chen, Jiayuan Mao

Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as the video-action generalization gap. In this pape…

View free PDFSource page
arxivcs.ROcs.AI2026-07-08

EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data

Baoyu Li, Xinchen Yin, Mengying Lin, Yixin Zhang, Danfei Xu

Egocentric human data offers scalable supervision for robot manipulation. However, behavior cloning entangles transferable content like objects, scenes, and task semantics, with non-transferable factors like human morphology, head motion, and behavioral style. We study whether Wo…

View free PDFSource page
arxivcs.RO2026-06-29

WARP: Whole-Body Retargeting for Learning from Offline Human Demonstrations

Zhenyang Chen, Chuizheng Kong, Chuye Zhang, Yuanshao Yang, Lawrence Y. Zhu, Shreyas Kousik, et al.

Direct transfer from human demonstration to learnable robot action is a crucial step towards scalable whole-body mobile manipulation. While human data scales better than mobile teleoperation, it requires overcoming significant embodiment gaps. Existing retargeting methods yield i…

View free PDFSource page
arxivcs.ROcs.AI2026-06-27

Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning

Shuo Cheng, Chuye Zhang, Alfred Cueva, Caelan Garrett, Ajay Mandlekar, Danfei Xu

Human videos are a scalable source of supervision for robot manipulation, as they are abundant and naturally capture rich object interactions. However, transferring human demonstrations to robots remains challenging due to embodiment mismatch, scene variation, and robot-specific…

View free PDFSource page
arxivcs.RO2026-06-26

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

Nadun Ranawaka, Josiah Wong, Wei-Lin Pai, Wei-Teng Chu, Tianyuan Dai, Masoud Moghani, et al.

Training and evaluating robot policies in the real world is costly and difficult to scale. We introduce SimFoundry, a modular and automated system for zero-shot real-to-sim scene construction from a video. SimFoundry generates sim-ready digital twins and supports object, scene, a…

View free PDFSource page