CORTEXA
← Browse

Duong The Do

1 paper indexed

arxivcs.NI2026-07-18

Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty

Jiaming Cheng, Duong The Do, Duong Tung Nguyen

KV cache memory is a primary bottleneck in modern LLM serving systems deployed on GPU clusters. A fundamental challenge is that the KV cache must be reserved upon request arrival, while the output token length remains unknown until generation completes. Under-reservation triggers…

View free PDFSource page