CORTEXA
← Browse
arxiveess.SY2026-07-08

Reachability-Preserving Bellman Operator for the Discounted Reach-Cost Value Function: Uniting Hamilton-Jacobi Reachability and Reinforcement Learning

Isabelle El-Hajj, Prashant Solanki, Jasper van Beers, Coen de Visser, Erik-Jan van Kampen

Hamilton-Jacobi (HJ) reachability provides rigorous safety and reachability guarantees for continuous-time dynamical systems, but its numerical solution suffers from the curse of dimensionality. Deep reinforcement learning (DRL), by contrast, offers scalable sample-based methods. However, RL is typically built around additive cumulative rewards; whereas, reachability objectives are inherently non-additive. This mismatch makes a direct bridge between HJ reachability and RL nontrivial. Recent discounted formulations have either introduced contraction by altering the original reachability semantics, or preserved exact semantics on the HJ side without a corresponding Bellman fixed-point characterization. In this paper, we close this gap by building on a semantics-preserving discounted reach-based value function and deriving a non-additive Bellman operator whose unique fixed point exactly matches the value function in the HJ formulation. We prove that discounting makes this operator contractive, yielding existence, uniqueness, and convergence of value iteration. Furthermore, we establish the equivalence between the HJ and Bellman characterizations, and show that RL can be interpreted as a sample-based approximation scheme for the same fixed-point equation. This yields a principled and semantically exact connection between HJ reachability and RL, enabling learning-based methods to approximate reachability value functions while preserving their safety-critical meaning. As a result, the proposed framework opens the door to scalable, data-driven computation of reachable sets and safety certificates in high-dimensional systems. Numerical experiments demonstrate close agreement with HJ solutions, confirm preservation of reachability semantics via alignment of zero level sets, and support the interpretation of reinforcement learning as a sample-based solver of the proposed Bellman operator.

View free PDFSource page

Related papers

arxiveess.SY2026-07-19

An Update to the Level Set Theorems in Hamilton-Jacobi Reachability Analysis

Dylan Hirsch, William McEneaney, Jaime Fisac, Claire Tomlin, Sylvia Herbert

Hamilton-Jacobi Reachability (HJR) is an important framework for controlling safety-critical systems despite uncertainty. Its theoretical underpinnings are rooted in Hamilton-Jacobi Partial Differential Equations, which provide the value function used for controller synthesis. Th…

View free PDFSource page
arxiveess.SYq-bio.QM2026-07-15

Exact Decomposition of Adversarial Dual-Objective Value Functions, with Applications to Optimal Drug Dosing

Dylan Hirsch, William Sharpless, Sylvia Herbert

Hamilton-Jacobi Reachability (HJR) is a central framework in safe control theory. While HJR has traditionally focused on a few fundamental tasks, there is increasing interest in scaling to more complex objectives. Recent works have studied the exact decomposition of the value fun…

View free PDFSource page
arxiveess.SYphysics.space-ph2026-07-02

Reachability-Based Safe-Start Regions for Approach to a Tumbling Target with Rotating LOS Constraints

Omer Burak Iskender, Keck Voon Ling, Wee Seng Lim, Erick Lansard

This paper presents a reachability-aware guidance architecture for autonomous approach to a tumbling, uncooperative target under a rotating line-of-sight (LOS) docking corridor. The LOS admissible set rotates with the target body frame, producing time-varying polyhedral constrain…

View free PDFSource page
arxivmath.OCeess.SY2026-07-08

Improving greenhouse fruit-production control by integrating reinforcement learning into short-horizon model predictive control

Bart van Laatum, Salim Msaad, Eldert J. van Henten, Robert D. McAllister, Sjoerd Boersma

Greenhouse fruit-production control aims to maximize the economic performance (fruit revenue minus operating costs) while operating within system constraints under external weather disturbances. Control methods need to balance the delayed economic benefit of fruit yield with curr…

View free PDFSource page
arxivcs.ROcs.AIeess.SY2026-07-20

The Open Ant: A Robot Platform for Reinforcement Learning Research

Elena Sorina Lupu, Patrick Spieler, Khurram Javed, Kris De Asis, John D. Martin, Martha Steenstrup, et al.

Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology remains rooted in simulations. The predominance of simulations makes translating research to physical reality uncertain for both algorithms a…

View free PDFSource page