CORTEXA
← Browse

Hao Huang

2 papers indexed

arxivcs.LGcs.AI2026-07-01

Muon as a Residual Connection

Hao Huang

Muon has recently emerged as one of the most effective optimizers for training large neural networks, yet its empirical success has been explained from several different perspectives. In this paper, we propose a simple mechanistic interpretation: Muon can be understood as an impl…

View free PDFSource page
arxivcs.LGcs.AI2026-06-27

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration

Gradwell Dzikanyanga, Yanqi Pan, Weihao Yang, Donglei Wu, Wen Xia, Hao Huang

Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads. Low-bit KV-cache quantization reduces this cost, yet it severely degrade quality; particularly, one-bit quantization reduce…

View free PDFSource page