arxivcs.CV2026-07-05
Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment
Zixiang Zhou, Zhentao Yu, Yifeng Ma, Hongmei Wang, Wenqing Yu, Cong Wang, et al.
Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity consistency and model complex relationships among multiple subjects. In this paper, we propose Aura, a unified framework for hig…