arxivcs.LGcs.AI2026-06-29
Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
Chao Han, Haozhe Hu, Xiaoyu Shen
Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive sparsification can trigger rapid performance degradation beyond an essential sparsity boundary. This work asks \emph{whether combining these two m…