CORTEXA
← Browse

Weiwei Lv

1 paper indexed

arxivcs.DCcs.AIcs.LG2026-07-03

SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference

Huajun Bai, Weiwei Lv, Huichuan Zheng, Youyou Lu, Jiwu Shu

LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: the model reasons, emits a tool call, then idles the GPU until the result returns. This wait consumes 16-37% of wall time in our wor…

View free PDFSource page