arxivcs.ARcs.AIcs.DC2026-07-19
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts
Pratyush Dhingra, Pramit Kumar Pal, Janardhan Rao Doppa, Partha Pratim Pande
Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained by three fundamental bottlenecks. These encompass the massive memory bandwidth required to fetch non…