arxivcs.LGeess.SY2026-07-24
Variance-Reduced Q-Learning over Static and Time-Varying Networks
Sreejeet Maity, Feng Zhu, Aritra Mitra, Robert W. Heath
We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, w…