arxivcs.CV2026-07-23
Out of Sight, Still in Mind: Token Compression for Omni-LLMs
Suho Yoo, Youngjoon Jang, Hyebin Cho, Joon Son Chung
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but the cost of the three streams is highly unbalanced: visual tokens account for the vast majority of…