Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework
Wanxin Jin, Zhaoran Wang, Zhuoran Yang, Shaoshuai Mou
摘要
This paper develops a Pontryagin Differentiable Programming (PDP) methodology, which establishes a unified framework to solve a broad class of learning and control tasks. The PDP distinguishes from existing methods by two novel techniques: first, we differentiate through Pontryagin's Maximum Principle, and this allows to obtain the analytical derivative of a trajectory with respect to tunable parameters within an optimal control system, enabling end-to-end learning of dynamics, policies, or/and control objective functions; and second, we propose an auxiliary control system in the backward pass of the PDP framework, and the output of this auxiliary control system is the analytical derivative of the original system's trajectory with respect to the parameters, which can be iteratively solved using standard control tools. We investigate three learning modes of the PDP: inverse reinforcement learning, system identification, and control/planning. We demonstrate the capability of the PDP in each learning mode on different high-dimensional systems, including multi-link robot arm, 6-DoF maneuvering quadrotor, and 6-DoF rocket powered landing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Safe Pontryagin Differentiable ProgrammingWanxin Jin, Shaoshuai Mou, George J. PappasNeurIPS 2021 · 被引用 65 次
- End-to-End Learning and Intervention in GamesJiayang Li, Jing Yu, Yu Marco Nie, Zhaoran WangNeurIPS 2020 · 被引用 48 次
- Adversarially robust learning for security-constrained optimal power flowPriya L. Donti, Aayushya Agarwal, Neeraj Vijay Bedmutha, Larry T. Pileggi 等NeurIPS 2021 · 被引用 25 次
- Revisiting Implicit Differentiation for Learning Problems in Optimal ControlMing Xu, Timothy L. Molloy, Stephen GouldNeurIPS 2023 · 被引用 16 次
- Neural Time-Reversed Generalized Riccati EquationAlessandro Betti, Michele Casoni, Marco Gori, Simone Marullo 等AAAI 2024 · 被引用 3 次
相关 Paper
- DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation LearningWeikang Wan, Ziyu Wang, Yufei Wang, Zackory Erickson 等NeurIPS 2024
- Unlocking Efficient Vehicle Dynamics Modeling via Analytic World ModelsAsen Nachkov, Danda Pani Paudel, Jan-Nico Zaech, Davide Scaramuzza 等AAAI 2026 · 被引用 2 次
- PODS: Policy Optimization via Differentiable SimulationMiguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev 等ICML 2021 · 被引用 65 次
- DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit DifferentiationShuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni 等ICML 2025
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 被引用 47 次
