arxivcs.AIcs.LG2026-07-03
MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference
Jingquan Chen, Jinghua Piao, Jie Feng, Shaogang Hu, Yong Li
Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these applications often incur high inference cost. We present MiniCache, a reusable program caching framework that transforms Program-of-Th…