CORTEXA
← Browse

Wenbo Hu

2 papers indexed

arxivcs.CV2026-07-04

G$^2$TAM: Geometry Grounded Track Anything Model

Chenming Zhu, Peizhou Cao, Jingli Lin, Wenbo Hu, Yunlong Ran, Jiangmiao Pang, et al.

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explicit object appearance memory banks for instance tracking, yet…

View free PDFSource page
arxivcs.CV2026-06-25

RayPE: Ray-Space Positional Encoding for 3D-Aware Video Generation

Minghao Yin, Jiahao Lu, Wenbo Hu, Wang Zhao, Shan Ying, Kai Han

Modern video diffusion transformers position their tokens through RoPE on the (u,v,t) axes -- a description of the camera's sampling grid that says nothing about the 3D structure of the scene. We observe that the geometric relation between two camera rays is captured by the Pluck…

View free PDFSource page