arxivcs.LGcs.CL2026-07-09
Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models
Bishmoy Paul, Youngmin Yi, Hoeseok Yang
Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation sparsification and token-level conditional routing. We first propose Sensiti…