arxivcs.CVcs.CLcs.LG2026-07-03
Pathways of Visual Information Flow in Vision-Language Models
Israfel Salazar, Stella Frank, Dan Oneata, Desmond Elliott, Constanza Fierro
We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we find that models rely on two distinct pathways to solve visual tasks: A direct pathway, where visual information is retained in image…