arxivcs.CV2026-07-15
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Zhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimized for pixel-level reconstruction, which can limit the semantic and spatio-temporal structure captured by their latents. Meanwh…