arxivcs.CVcs.AIcs.MMcs.SD2026-07-16
SceneBind: Binding What and Where Across Vision, Audio and Language
Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial struc…