CORTEXA
← Browse

Yuxuan Yang

1 paper indexed

arxivcs.LGcs.AI2026-06-27

HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression

Yuxuan Yang, Feiyang Ren, Bowen Zeng, Dalin Zhang, Jinpeng Chen, Gang Chen, et al.

Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top-$p$ nucleus sampling) offer superior accuracy by dynamically fluctuating memory budgets, yet modern inference engines (e.g., vLLM) demand rigid, static memory patterns to leve…

View free PDFSource page