arxivcs.CVcs.LG2026-06-30
Steal the Patch Size: Adversarially Manipulate Vision-Language Models
Kai Hu, Akash Bharadwaj, Weichen Yu, Matt Fredrikson
We present a black-box model-stealing attack that recovers private vision-tokenizer configurations of deployed vision-language models (VLMs), including the visual patch size and input preprocessing pipeline. The key idea is a task-level side channel induced by ViT-style patchific…