CORTEXA
← Browse

Xuqing Yang

1 paper indexed

arxivcs.AI2026-07-02

Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling

Xuqing Yang, Yi Yuan, Shanzhe Lei, Xuhong Wang

Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing RL reward designs typically prioritize response correctness, neglecting to incentivize models to expr…

View free PDFSource page