arxivcs.AI2026-07-15
Reward-Free Evolving Agents via Pairwise Validator
Minghao Liu, Yu Wang, Jiayun Wang, Wei Wei
A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the change based on a per-iteration quality signal. Designing that signal is often the costly part of the project: a reliable scalar reward requi…