arxivcs.AIcs.CVcs.LGcs.MM2026-07-18
From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence
Nadine Chang, Maying Shen, Shizhe Diao, Jialiang Wang, Jingde Chen, Thomas Breuel, et al.
We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene. A global semantic codebook unifies these into a…