arxiveess.SYcs.LG2026-07-14
Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning
Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tuned to shape the transition dynamics and costs exp…