CORTEXA
← Browse

Minghao Yin

1 paper indexed

arxivcs.CV2026-06-25

RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation

Minghao Yin, Jiahao Lu, Wenbo Hu, Wang Zhao, Shan Ying, Kai Han

Modern video diffusion transformers position their tokens through RoPE on the (u,v,t) axes -- a description of the camera's sampling grid that says nothing about the 3D structure of the scene. We observe that the geometric relation between two camera rays is captured by the Pluck…

View free PDFSource page