CORTEXA
← Browse

Duong Tung Nguyen

2 papers indexed

arxivcs.NI2026-07-18

Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty

Jiaming Cheng, Duong The Do, Duong Tung Nguyen

KV cache memory is a primary bottleneck in modern LLM serving systems deployed on GPU clusters. A fundamental challenge is that the KV cache must be reserved upon request arrival, while the output token length remains unknown until generation completes. Under-reservation triggers…

View free PDFSource page
arxiveess.SY2026-07-14

Scenario-Free Uncertainty-Aware DLMP-Based Bilevel Coordination of EV Charging and Reactive Power Support in Distribution Networks

Arash Baharvandi, Duong Tung Nguyen

This paper develops a scenario-free uncertainty-aware bilevel optimization framework for coordinated electric vehicle (EV) charging and reactive power support in distribution networks using distribution locational marginal prices (DLMPs). The upper-level EV aggregator jointly sch…

View free PDFSource page