arxivcs.CVcs.AIcs.ARcs.DCcs.LG2026-07-01
Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers
Aravind Pradeep, Samira Nazari, Mahdi Taheri, Christian Herglotz
Vision Transformers achieve strong image classification accuracy but process all image regions with nearly the same computation, even when many regions are redundant or uninformative. Recent adaptive inference methods reduce this cost by selectively compressing tokens or terminat…