Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework
Wanxin Jin, Zhaoran Wang, Zhuoran Yang, Shaoshuai Mou
Abstract
This paper develops a Pontryagin Differentiable Programming (PDP) methodology, which establishes a unified framework to solve a broad class of learning and control tasks. The PDP distinguishes from existing methods by two novel techniques: first, we differentiate through Pontryagin's Maximum Principle, and this allows to obtain the analytical derivative of a trajectory with respect to tunable parameters within an optimal control system, enabling end-to-end learning of dynamics, policies, or/and control objective functions; and second, we propose an auxiliary control system in the backward pass of the PDP framework, and the output of this auxiliary control system is the analytical derivative of the original system's trajectory with respect to the parameters, which can be iteratively solved using standard control tools. We investigate three learning modes of the PDP: inverse reinforcement learning, system identification, and control/planning. We demonstrate the capability of the PDP in each learning mode on different high-dimensional systems, including multi-link robot arm, 6-DoF maneuvering quadrotor, and 6-DoF rocket powered landing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c170a1c3-6d40-4a2d-9e9d-498b43e88967Cited by top-tier papers10
- Safe Pontryagin Differentiable ProgrammingWanxin Jin, Shaoshuai Mou, George J. PappasNeurIPS 2021 · 65 citations
- End-to-End Learning and Intervention in GamesJiayang Li, Jing Yu, Yu Marco Nie, Zhaoran WangNeurIPS 2020 · 48 citations
- Adversarially robust learning for security-constrained optimal power flowPriya L. Donti, Aayushya Agarwal, Neeraj Vijay Bedmutha, Larry T. Pileggi et al.NeurIPS 2021 · 25 citations
- Revisiting Implicit Differentiation for Learning Problems in Optimal ControlMing Xu, Timothy L. Molloy, Stephen GouldNeurIPS 2023 · 16 citations
- Neural Time-Reversed Generalized Riccati EquationAlessandro Betti, Michele Casoni, Marco Gori, Simone Marullo et al.AAAI 2024 · 3 citations
Related papers
- DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation LearningWeikang Wan, Ziyu Wang, Yufei Wang, Zackory Erickson et al.NeurIPS 2024
- Unlocking Efficient Vehicle Dynamics Modeling via Analytic World ModelsAsen Nachkov, Danda Pani Paudel, Jan-Nico Zaech, Davide Scaramuzza et al.AAAI 2026 · 2 citations
- PODS: Policy Optimization via Differentiable SimulationMiguel Zamora, Momchil Peychev, Sehoon Ha, Martin T. Vechev et al.ICML 2021 · 65 citations
- DiLQR: Differentiable Iterative Linear Quadratic Regulator via Implicit DifferentiationShuyuan Wang, Philip D. Loewen, Michael G. Forbes, R. Bhushan Gopaluni et al.ICML 2025
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 47 citations
