CORTEXA
← Browse

Zhuoyi Lin

1 paper indexed

arxivcs.LG2026-07-05

On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models

Chee Heng Tan, Zhuoyi Lin, Mehul Motani, Wee Sun Lee

In this paper, we consider the setting where large language models (LLMs) are trained using reinforcement learning (RL) to simultaneously improve reasoning accuracy and verbalize its confidence. Our reward scheme uses two functions for rewarding confidence verbalized by the LLM:…

View free PDFSource page