arxivcs.CV2026-07-19
Searching for Task-Specific Vision Paths: Evolutionary Block Pruning Across Vision-Language Models
Vision-language models normally execute the same complete vision encoder for every question, even when OCR, counting, object, attribute, and spatial queries may not require identical computation. We study whether fixed-budget combinations of vision blocks can be skipped without f…