CORTEXA
← Browse

Zijun Cui

3 papers indexed

arxivcs.CVcs.AIcs.MMcs.SD2026-07-16

SceneBind: Binding What and Where Across Vision, Audio and Language

Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial struc…

View free PDFSource page