arxivcs.LG2026-07-11
Empowering Long-form Omni-modal Understanding with Robust Audio Perception
Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie
Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues. To bridgethis gap,…