CORTEXA
← Browse

Bishmoy Paul

1 paper indexed

arxivcs.LGcs.CL2026-07-09

Sensitivity-Aware Thresholding and Token Routing for Activation Sparsification in Large Language Models

Bishmoy Paul, Youngmin Yi, Hoeseok Yang

Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality. We study this problem through multilayer perceptron (MLP) activation sparsification and token-level conditional routing. We first propose Sensiti…

View free PDFSource page