arxivcs.LGcs.ARcs.OS2026-07-16
PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference
Hyunwoo Oh, Suyeon Jang, Hanning Chen, KyungIn Nam, Sanggeon Yun, Ryozo Masukawa, et al.
CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs. We present PolyQ, a CPU-oriented compiler/quantization…