arxivcs.LGcs.AIcs.GTmath.CO2026-07-09
AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision
Brent Kong, Tejas Ram, Tony Yue Yu
AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necessarily imply perfect play. We study this gap in two oracle-evaluable domains with contrasting structure: Connect Four, a solved partisan game…