arxivcs.SEcs.LGcs.PL2026-07-15
Quantize with Confidence? An Empirical Study of Quantization for Code Generation
Saima Afrin, Md. Zahidul Haque, Antonio Mastropaolo
The growing adoption of local inference frameworks such as Ollama has made it increasingly common for developers to run large code models on laptops and other resource-constrained hardware. In these settings, post-training quantization is essential for reducing memory footprint a…