arxivcs.NEcs.AIcs.CVcs.LG2026-07-07
Do You Remember? Toward Memory-Centric Multimodal AI
Xuguang Yu, Weigang Zheng, Minyue Yu
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen visual encoder, produce a one-shot text output, and discard internal representations. We present DoYouRemember, a three-stage archi…