Solving Minimum-Cost Reach Avoid using Reinforcement Learning
Oswin So, Cheng Ge, Chuchu Fan
Abstract
Current reinforcement-learning methods are unable to directly learn policies that solve the minimum cost reach-avoid problem to minimize cumulative costs subject to the constraints of reaching the goal and avoiding unsafe states, as the structure of this new optimization problem is incompatible with current methods. Instead, a surrogate problem is solved where all objectives are combined with a weighted sum. However, this surrogate objective results in suboptimal policies that do not directly minimize the cumulative cost. In this work, we propose RC-PPO, a reinforcement-learning-based method for solving the minimum-cost reach-avoid problem by using connections to Hamilton-Jacobi reachability. Empirical results demonstrate that RC-PPO learns policies with comparable goal-reaching rates to while achieving up to 57% lower cumulative costs compared to existing methods on a suite of minimum-cost reach-avoid benchmarks on the Mujoco simulator. The project page can be found at https://oswinso.xyz/rcppo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext baa52cf6-1505-43ec-a225-160586a6bfa6Cited by top-tier papers7
- Dual-Objective Reinforcement Learning with Novel Hamilton-Jacobi-Bellman FormulationsWilliam Sharpless, Dylan Hirsch, Sander Tonkens, Nikhil Uday Shinde et al.ICLR 2026 · 12 citations
- Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement LearningOswin So, Eric Yang Yu, Songyuan Zhang, Matthew Cleaveland et al.ICLR 2026 · 1 citation
- Safe Continuous-time Multi-Agent Reinforcement Learning via Epigraph FormXuefeng Wang, Lei Zhang, Henglin Pu, Husheng Li et al.ICLR 2026 · 1 citation
- HMARL-CBF - Hierarchical Multi-Agent Reinforcement Learning with Control Barrier Functions for Safety-Critical Autonomous SystemsH. M. Sabbir Ahmad, Ehsan Sabouni, Alexander Wasilkoff, Param Budhraja et al.NeurIPS 2025
- A Physics-Informed Machine Learning Framework for Safe and Optimal Control of Autonomous SystemsManan Tayal, Aditya Singh, Shishir Kolathaya, Somil BansalICML 2025
Builds on11
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Contrastive Learning as Goal-Conditioned Reinforcement LearningBenjamin Eysenbach, Tianjun Zhang, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 331 citations
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu et al.ICLR 2021 · 222 citations
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner et al.NeurIPS 2021 · 177 citations
- Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal GenerationSuraj Nair, Chelsea FinnICLR 2020 · 152 citations
Related papers
- Stochastic Minimum-Cost Reach-Avoid Reinforcement LearningJingduo Pan, Taoran Wu, Yiling Xue, Bai XueICML 2026
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 58 citations
- Iterative Reachability Estimation for Safe Reinforcement LearningMilan Ganai, Zheng Gong, Chenning Yu, Sylvia L. Herbert et al.NeurIPS 2023 · 56 citations
- Reachability Constrained Reinforcement LearningDongjie Yu, Haitong Ma, Sheng-bo Li, Jianyu ChenICML 2022 · 90 citations
- Constrained Update Projection Approach to Safe Policy OptimizationLong Yang, Jiaming Ji, Juntao Dai, Linrui Zhang et al.NeurIPS 2022 · 95 citations
