arxivcs.NI2026-07-09
MORES: Mobile Reasoning-as-a-Service via Distributed LLM Inference-Time Scaling
Guanchen Liu, Hongyang Du, Kaibin Huang
Inference-time scaling has emerged as an effective approach for enhancing the capabilities of Large Language Models (LLMs), addressing the growing demand for stronger reasoning without increasing model size. This novel form of LLM scaling comprises two representative approaches:…