CORTEXA
← Browse

Sadra Saremi

1 paper indexed

arxivcs.LG2026-07-04

AdaptiveSD A Stability-Aware, Runtime-Adaptive Speculative Decoding Framework with Multi-Policy Orchestration for CPU-Constrained LLM Inference

Sadra Saremi

With the rise of small quantized GGUF-based language models and their increasing use for on-device inference tasks, we have seen the growing need for an approach capable of reliably delivering these models at scale even under severe memory bandwidth constraints such as those impo…

View free PDFSource page