CORTEXA
← Browse

Yao Lyu

1 paper indexed

arxivcs.LG2026-07-20

Distributional Soft Bellman Operator under the Cramér Geometry

Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li

Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regulari…

View free PDFSource page