CORTEXA
← Browse

Wangbo Zhao

1 paper indexed

arxivcs.LG2026-06-26

FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

Fan Mo, Yuxuan Han, Geng Zhang, Wangbo Zhao, Yang You

Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of storing and serving all experts, and the available…

View free PDFSource page