arxivcs.AIcs.CLcs.CVcs.MM2026-07-14
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
Kaiwen Zheng, Junchen Fu, Wenhao Deng, Hu Han, Joemon M. Jose, Xuri Ge
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improve…