CORTEXA
← Browse
arxiveess.SYcs.LG2026-07-14

Environment Parameter Gradient Theorem for Policy-Environment Co-Design in Reinforcement Learning

Amber Srivastava

Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tuned to shape the transition dynamics and costs experienced by the agent. This motivates jointly optimizing both the policy and the environment design parameters. To this end, we establish an Environment Parameter Gradient Theorem -- a formal expression for the gradient of the value function with respect to environment parameters. The key theoretical device is a generalized action-value function $Q_{π,ξ}(s,a,ζ)$, which comprises two copies of the environment parameters: $ζ$ governs the cost and transition dynamics at the current state--action pair, while $ξ$ governs the future rollouts. This decoupling yields a tractable closed-form gradient expression and is essential to the theorem's derivation. Building on this result, we develop a model-free algorithm that simultaneously learns the optimal policy and the environment parameters. We demonstrate the efficacy of our framework on a UAV network design problem, where the optimal UAV placement (environment parameters) and communication routes (governed by the policy) are learned jointly to minimize the total communication cost in the network.

View free PDFSource page

Related papers

arxivcs.ROcs.LGeess.SY2026-07-15

Flow-aware Optimal Navigation in Unsteady Flows through Reinforcement Learning

Andrea Maria Braghin, Nicolò Botteghi, Matteo Tomasetto, Andrea Manzoni, Gabriele Cazzulani

Autonomous robotic navigation in nonstationary time-varying fluid flows remains a fundamental challenge due to partial observability and the unpredictability of realistic environments. While classical optimal control frameworks employed in robotics require unrealistic a-priori gl…

View free PDFSource page
arxivcs.LGeess.SY2026-07-01

Wind-Aware Reinforcement Learning Control of a Small Quadrotor Using Learned Onboard Wind Estimation in Simulated Atmospheric Turbulence

Abdullah Al Tasim, Wei Sun

Small multirotor aircraft are increasingly tasked with operations in the atmospheric boundary layer, where turbulent winds comparable to the vehicle's airspeed degrade trajectory tracking and can defeat conventional feedback control. This work illustrates a two-stage learning pip…

View free PDFSource page
arxivcs.LGcs.AIeess.SY2026-06-30

Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority

Philip-Roman Adam, Stefanie Schmidtner

Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles. Existing reinforcement-learning (RL) approaches to TSP typically encode transit-aware f…

View free PDFSource page
arxivcs.ROcs.LGcs.MAeess.SY2026-07-10

Runtime Safety Filtering for Learned Small UAS Separation Policies under GNSS Degradation

Alex Zongo, Peng Wei

Learning-based separation assurance for small Unmanned Aircraft Systems (sUAS) achieves near-zero collision rates in simulation, but assumes accurate position and velocity information from Global Navigation Satellite Systems (GNSS). This assumption fails in urban environments, wh…

View free PDFSource page
arxivcs.LGeess.SY2026-07-06

Federated Physics-Grounded Reinforcement Learning for Distributed Stability Control in Smart Grids

Omar Al-Refai, Ibrahim Shahbaz, Adam Ali Husseinat, Eman Hammad

Transient stability control in smart grids requires rapid post-fault damping of generator frequency and rotor angle deviations to prevent cascading failures. This paper proposes FedPPO-PG, a Federated Multi-Agent Proximal Policy Optimization framework with Physics-Grounded neighb…

View free PDFSource page