DiffOP: Reinforcement Learning of Optimization-Based Control Policies via Implicit Policy Gradients
Yuexin Bian, Jie Feng, Yuanyuan Shi
Abstract
Real-world control systems require policies that are not only high-performing but also interpretable and robust. A promising direction toward this goal is model-based control, which learns system dynamics and cost functions from historical data and then uses these models to inform decision-making. Building on this paradigm, we introduce DiffOP, a novel framework for learning optimization-based control policies defined implicitly through optimization control problems. Without relying on value function approximation, DiffOP jointly learns the cost and dynamics models and directly optimizes the actual control costs using policy gradients. To enable this, we derive analytical policy gradients by applying implicit differentiation to the underlying optimization problem and integrating it with the standard policy gradient framework. Under standard regularity conditions, we establish that DiffOP converges to an ϵ-stationary point within O(ϵ -1 ) iterations. We demonstrate the effectiveness of DiffOP through experiments on nonlinear control tasks and power system voltage control with constraints. The code is available at https://github.com/alwaysbyx/DiffOP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 343 citations
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 162 citations
- Pontryagin Differentiable Programming: An End-to-End Learning and Control FrameworkWanxin Jin, Zhaoran Wang, Zhuoran Yang, Shaoshuai MouNeurIPS 2020 · 133 citations
- A Fully First-Order Method for Stochastic Bilevel OptimizationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICML 2023 · 123 citations
- The Curse of Unrolling: Rate of Differentiating Through OptimizationDamien Scieur, Gauthier Gidel, Quentin Bertrand, Fabian PedregosaNeurIPS 2022 · 20 citations
Related papers
- DiffTORI: Differentiable Trajectory Optimization for Deep Reinforcement and Imitation LearningWeikang Wan, Ziyu Wang, Yufei Wang, Zackory Erickson et al.NeurIPS 2024
- Control-Oriented Model-Based Reinforcement Learning with Implicit DifferentiationEvgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc BaconAAAI 2022 · 47 citations
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin et al.ICLR 2023
- Gradient Information Matters in Policy Optimization by Back-propagating through ModelChongchong Li, Yue Wang, Wei Chen, Yuting Liu et al.ICLR 2022 · 10 citations
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy OptimizationPierluca D'Oro, Wojciech JaskowskiNeurIPS 2020 · 33 citations
