Safe Pontryagin Differentiable Programming
Wanxin Jin, Shaoshuai Mou, George J. Pappas
摘要
We propose a Safe Pontryagin Differentiable Programming (Safe PDP) methodology, which establishes a theoretical and algorithmic framework to solve a broad class of safety-critical learning and control tasks-problems that require the guarantee of safety constraint satisfaction at any stage of the learning and control progress. In the spirit of interior-point methods, Safe PDP handles different types of system constraints on states and inputs by incorporating them into the cost or loss through barrier functions. We prove three fundamentals of the proposed Safe PDP: first, both the solution and its gradient in the backward pass can be approximated by solving their more efficient unconstrained counterparts; second, the approximation for both the solution and its gradient can be controlled for arbitrary accuracy by a barrier parameter; and third, importantly, all intermediate results throughout the approximation and optimization strictly respect the constraints, thus guaranteeing safety throughout the entire learning and control process. We demonstrate the capabilities of Safe PDP in solving various safety-critical tasks, including safe policy optimization, safe motion planning, and learning MPCs from demonstrations, on different challenging systems such as 6-DoF maneuvering quadrotor and 6-DoF rocket powered landing. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Enforcing Hard Constraints with Soft Barriers: Safe Reinforcement Learning in Unknown Stochastic EnvironmentsYixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang 等ICML 2023 · 被引用 81 次
- Revisiting Implicit Differentiation for Learning Problems in Optimal ControlMing Xu, Timothy L. Molloy, Stephen GouldNeurIPS 2023 · 被引用 16 次
- DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit DifferentiationShuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni 等ICML 2025
- DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation LearningWeikang Wan, Ziyu Wang, Yufei Wang, Zackory Erickson 等NeurIPS 2024
它引用的顶会 Paper6
- What is Local Optimality in Nonconvex-Nonconcave Minimax Optimization?Chi Jin, Praneeth Netrapalli, Michael I. JordanICML 2020 · 被引用 381 次
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- IPO: Interior-Point Policy Optimization under ConstraintsYongshuai Liu, Jiaxin Ding, Xin LiuAAAI 2020 · 被引用 231 次
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 被引用 190 次
- Pontryagin Differentiable Programming: An End-to-End Learning and Control FrameworkWanxin Jin, Zhaoran Wang, Zhuoran Yang, Shaoshuai MouNeurIPS 2020 · 被引用 133 次
相关 Paper
- Safe DNN-type Controller Synthesis for Nonlinear Systems via Meta Reinforcement LearningHanrui Zhao, Xia Zeng, Niuniu Qi, Zhengfeng Yang 等DAC 2023 · 被引用 4 次
- Infinite-Horizon Differentiable Model Predictive ControlSebastian East, Marco Gallieri, Jonathan Masci, Jan Koutník 等ICLR 2020 · 被引用 39 次
- Proactive Constrained Policy Optimization with Preemptive PenaltyNing Yang, Pengyu Wang, Guoqing Liu, Haifeng Zhang 等AAAI 2026 · 被引用 1 次
- Hybrid Controller Synthesis for Nonlinear Systems Subject to Reach-Avoid ConstraintsZhengfeng Yang, Li Zhang, Xia Zeng, Xiaochao Tang 等CAV 2023 · 被引用 6 次
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
