arxivcs.AR2026-07-25
Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch
En-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng, Shih-Hao Hung, H. T. Kung
Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kernels using static token-count buckets, ignoring the per-expert routing distribution that determines t…