arxivcs.CLcs.LG2026-07-16
VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
Shahrzad Esmat, Dhawal Shah, Ali Jannesari
The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict low-scorin…