CORTEXA
← Browse

Wai Tsang Keung

1 paper indexed

arxivcs.LGstat.ML2026-07-22

Efficient Clustering with Provable Guardrails for LLM Inference at Scale

Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara, Roman Wang, Ali Dashti, Francesc Moreno-Noguer

Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models. A natural fix is to cluster the inputs and call the LLM only on cluster representatives, letting other members inherit the output -- but this is only…

View free PDFSource page