CORTEXA
← Browse
arxivmath.OCcs.LGcs.MAmath.PR2026-07-01

Mean Field Reinforcement Learning

René Carmona, Mathieu Laurière

This monograph provides an introduction to mean field reinforcement learning through the lens of Markov decision processes arising from large-population stochastic control with mean field interactions and common noise. Starting from the connection between multi-agent reinforcement learning and mean field control, it develops the probabilistic, mathematical, and control-theoretic framework needed to formulate representative-agent learning problems, analyze their relationship with finite-population systems, and study both general and linear-quadratic models. The presentation includes dynamic programming principles, propagation-of-chaos limits, and theoretical analyses of tabular Q-learning and policy-gradient methods. It also discusses numerical implementations, including tabular schemes and deep reinforcement learning methods such as deep deterministic policy gradient. The goal is to give readers a coherent bridge between mean field control theory and reinforcement learning methodology, emphasizing the mathematical structure of the problems and the design of tractable learning approaches for large stochastic populations.

View free PDFSource page

Related papers

arxivcs.LGcs.MAmath.OC2026-07-06

Deep Reinforcement Learning for Dynamic Battery Management of Autonomous Order Pickers

Taniya Shaji, Abhay Sobhanan, Christof Defryn

Battery charging of Autonomous Mobile Robots (AMRs) in warehouses is a critical operational challenge that heavily impacts both order processing times and throughput. In this study, we address the dynamic AMR charging problem under stochastic order arrivals, where robots must lea…

View free PDFSource page
arxivcs.GTcs.LGcs.MAmath.DSmath.OC2026-07-13

Paradoxes of Game Theoretic Equilibria and Price of Anarchy

Georgios Piliouras, Ian Gemp, Siqi Liu, Luke Marris

For decades, static solution concepts (Nash, Correlated, and Coarse Correlated Equilibria) and the Price of Anarchy (PoA) have formed the bedrock of algorithmic game theory, with no-regret learning proving fast convergence to such game-theoretic equilibria. We show that reducing…

View free PDFSource page
arxivmath.OCcs.LGmath.PRstat.ML2026-06-30

Homogenization of $\ell_2$-Adversarial Training in High-Dimensions: Exact Dynamics under Stochastic Gradient Descent

Fabrizzio Sabelli

We develop a framework for analyzing the learning dynamics of $\ell_2$-adversarial training of single-index models on Gaussian mixtures in the high-dimensional limit under streaming stochastic gradient descent (SGD). We derive deterministic equivalents for a broad class of statis…

View free PDFSource page
arxivmath.OCcs.LGcs.MAeess.SY2026-07-10

Control Laguerre Tessellation: Semi-discrete Optimal Transport Over Control Systems

Ripon C. Sarker, Abhishek Halder

We study the optimal transport of optimally controlled agents from a compactly supported absolutely continuous source to a discrete target measure. The ground cost for the transport is induced by the optimal cost of the agents' motion. When this ground cost satisfies the twist co…

View free PDFSource page