arxivcs.CVcs.CL2026-07-03
ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space
Peiming Li, Yifan Wang, Xiaotian Zhang, Zhiyuan Hu, Shiyu Li, Zheng Wei, et al.
Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduction. While explicit visual generation incurs prohibitive computational costs, existing latent approa…