arxivcs.CV2026-07-01
Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers
Cong Liu, Xiaofang Li, Simon X. Yang
Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial regularities of natural images, we ask whether spatial organization can be induced from data rather than explicitly injected. Un…