arxivcs.AIcs.CLcs.LG2026-06-29
When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models
Zhe Dong, Fang Qin, Manish Shah
Reasoning models spend test-time compute unevenly across instances, and a growing family of early-exit rules -- confidence thresholds, entropy monitors, answer-stability checks, and learned stoppers -- promises to reclaim the waste. These rules, however, are evaluated under heter…