CORTEXA
← Browse

Xiaohan Zhang

6 papers indexed

arxivcs.CV2026-07-15

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, et al.

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyfram…

View free PDFSource page
arxivcs.CVcs.SD2026-07-15

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

Xiaohan Zhang, Yuqing Wen, Junlin Chen, Yuqi Tang, Yiting He, Lizhuo Shao, et al.

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video…

View free PDFSource page
arxivcs.CV2026-07-08

GIRAF: Towards Generalizable Human Interactions with Articulated Objects

Xiaohan Zhang, Sebastian Starke, Alexander Winkler, Federica Bogo, Samir Aroudj, Yuting Ye

Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in robotics training and virtual agents. Existing models remain limited: some focus on simple activities with static objects, wh…

View free PDFSource page
arxivcs.GRcs.AI2026-07-04

CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation

Zhenyu Sun, Xiaohan Zhang, Qi Liu, Huan Wang

Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the constr…

View free PDFSource page
arxivcs.CV2026-07-03

FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers

Chen Chen, Yuanmin Huang, Zhenfei Zhang, Mi Zhang, Xiaohan Zhang, Yun Xiong, et al.

Multimodal diffusion transformers (MM-DiTs) have emerged as the prevalent backbone for modern text-to-image generation systems. However, they exhibit critical alignment vulnerabilities, systematically manifesting severe stereotype biases even under benign prompts. This poses a si…

View free PDFSource page
arxivcs.RO2026-06-30

Efficient Sim-to-Real Transfer of World-Action Models from Synthetic Priors

Zixing Wang, Kausik Sivakumar, Jinghuan Shang, Yafei Hu, Zhaoming Xie, Ran Gong, et al.

Bridging the sim-to-real gap is a core challenge in deploying learned manipulation policies. Sim-to-real learning is attractive because it can replace expensive real robot demonstrations with scalable synthetic data, yet world-action models have not previously been shown to trans…

View free PDFSource page