arxivcs.CVcs.AIcs.CL2026-07-01
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
Yixin Ji, Fanghua Ye, Juntao Li, Bo Zhao, Zexuan Qiu, Zhaopeng Tu, et al.
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing online methods either retain compact visual representations that lack semantic structure, or build…