arxivcs.LGcs.AI2026-07-12
MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference
Large language model (LLM) agents accumulate heterogeneous context, including system instructions, plans, user turns, retrieved documents, tool outputs, and intermediate reasoning, whose key-value (KV) cache can become a major memory bottleneck. Existing eviction policies general…