arxivcs.CLcs.LG2026-07-07
Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization
Kaishen Wang, Tong Zheng, Xuehao Cui, Ruibo Chen, Tianyi Xiong, Heng Huang
Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refi…