CORTEXA
← Browse
arxivcs.LGmath.OC2026-07-31

Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment

Yanwei Jia, Du Ouyang

This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the logit parameterization for the stochastic policy, we show that it converges almost surely to the optimal arm under an arbitrary constant learning rate. Furthermore, we derive the non-asymptotic regret upper bound when the constant learning rate is below a time-invariant threshold; and the regret bound has order $O(\log T)$. We improve the analysis in Lattimore (2026a) for the same SDE by constructing a novel Lyapunov function and demonstrate the transparency of analyzing policy gradient using the tools in SDEs. In addition, the same Lyapunov function is also helpful in analyzing the discrete-time policy gradient algorithm.

View free PDFSource page

Related papers

arxivstat.MLcs.LGmath.OC2026-07-08

Expressivity and Statistical Trade-offs in Diffusion Policy Learning

Viet Vu, Renyuan Xu, Jiacheng Zhang, Yufei Zhang

Diffusion-based policies have recently emerged as powerful policy parameterizations for reinforcement learning, representing state-conditioned action distributions as terminal laws of diffusion processes with parameterized drifts. This terminal-law representation has shown substa…

View free PDFSource page
arxivmath.OCcs.LG2026-07-05

Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks

Shokhrukh Ibragimov, Arnulf Jentzen

Gradient based optimization methods are nowadays the methods of choice for training deep neural networks (DNNs) in artificial intelligence (AI) systems. In practically relevant DNN training problems, one does usually not apply the standard gradient descent (GD) optimization metho…

View free PDFSource page
arxivcs.LGmath.OC2026-07-16

Regularity-Aware Stochastic MGDA with Adaptive Conflict-Avoidant Update Direction Control

Chentong Huang, Lisha Chen

Multi-objective learning (MOL) aims to optimize multiple objectives simultaneously. The multi-gradient descent algorithm (MGDA) is a workhorse that iteratively updates along a common descent or conflict-avoidant (CA) direction across objectives. In stochastic settings, however, t…

View free PDFSource page
arxivmath.OCcs.LGcs.MAmath.PR2026-07-01

Mean Field Reinforcement Learning

René Carmona, Mathieu Laurière

This monograph provides an introduction to mean field reinforcement learning through the lens of Markov decision processes arising from large-population stochastic control with mean field interactions and common noise. Starting from the connection between multi-agent reinforcemen…

View free PDFSource page
arxivcs.LGmath.OC2026-07-17

Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems

Matteo Tomasetto, Nicolò Botteghi, Gabriele Bruni, Andrea Manzoni

Reinforcement learning (RL) has recently emerged as a promising feedback control strategy for nonlinear and complex dynamical systems. However, RL algorithms are sample inefficient and require a large number of interaction with the environment to synthesize optimal control strate…

View free PDFSource page