CORTEXA
← Browse

Yuki M. Asano

2 papers indexed

arxivcs.LGcs.CV2026-07-14

AVQ-Attention: Adaptive Vector-Quantized Attention

Winfried van den dool, Patrick Forré, Amir Habibian, Yuki M. Asano, Max Welling

The $\mathcal{O}(N^2)$ complexity of attention over $N$ tokens remains a computational bottleneck in transformer models. Vector-Quantized (VQ) attention reduces this to $\mathcal{O}(MN)$ by representing keys with $M$ codewords, but applies uniform codebook capacity regardless of…

View free PDFSource page