arxivcs.LGcs.AIq-bio.BMstat.ML2026-07-01
Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization
Xuefeng Liu, Mingxuan Cao, Qinan Huang, Thomas Brettin, Rick Stevens, Le Cong
Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of training such reasoning remains a key open challenge. We study this problem in instruction-based molecular optimization, where answer-only supervi…