CORTEXA
← Browse

Yang You

2 papers indexed

arxivcs.LG2026-06-26

FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

Fan Mo, Yuxuan Han, Geng Zhang, Wangbo Zhao, Yang You

Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of storing and serving all experts, and the available…

View free PDFSource page