arxivcs.LGmath.OC2026-07-31
Convergence and Regret of the Policy Gradient for Multi-Armed Bandits in Diffusion Environment
This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al. (2020), Jia and Zhou (2022b). With the log…