CORTEXA
← Browse

Hongyao Tang

2 papers indexed

arxivcs.LG2026-06-28

The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning

Jing Liang, Hongyao Tang, Yi Ma, Yancheng He, Weixun Wang, Xiaoyang Li, et al.

Reinforcement learning (RL) has gained growing attention in large language model (LLM) post-training, yet RL training remains fragile and can suffer from instability or collapse. One vital cause is training-inference mismatch: LLM adopts separate inference and training engines fo…

View free PDFSource page
crossrefElectronics2024-12-20Cited by 1

Enhanced Sagger Crack Detection Integrating Deep Learning and Machine Vision

Tao Song, Ting Chen, Yuan Gong, Yulin Wang, Lu Ran, Jiale Chen, et al.

In recent years, target inspection has found extensive utilization within the industry, making it crucial to detect defects in industrial products to ensure quality. To address the challenges posed by large brightness differences, attached dirt, and complex backgrounds in saggers…

View free PDFSource page