arxivcs.LGcs.AI2026-07-06
Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation
Muhammad Zain Amin, Kibele Sebnem Yildirim
Reinforcement Learning is commonly used to train large language models using environmental feedback. In applied settings, the environment usually provides sparse or delayed feedback. This makes it difficult for the model to pinpoint which actions in its reasoning led to success o…