arxivcs.CV2026-07-15
CRISP: Pre-LLM Yet Text-Driven Visual Token Pruning for Efficient LVLM Inference
Xu Li, Yi Zheng, Mengyang Zhao, Yuxuan Liang, Zhe Liu, Rui Zhu, et al.
Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning methods either operate before the LLM using text-agnostic heuristics or prune inside the LLM at the c…