arxivcs.LGcs.AI2026-07-22
ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
Mahdi Heidari, Mohammad Mahdi Rahimi, Jaekyun Moon
The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths. Existing efficient attention methods usually reduce this bottleneck by either imposing sparsity, so that each query attends to only a small subset of key…