arxivcs.DCcs.LG2026-07-02
Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference
Wenchen Han, Gingfung Matthew Yeung, Marco Barletta, William Toner, Amory Hoste, Adam Barker
Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems. In disaggregated inference, these workloads require transferring large Key-Value (KV) caches across the network, where decoding canno…