Deep Conservative Policy Iteration
Nino Vieillard, Olivier Pietquin, Matthieu Geist
Abstract
Conservative Policy Iteration (CPI) is a founding algorithm of Approximate Dynamic Programming (ADP). Its core principle is to stabilize greediness through stochastic mixtures of consecutive policies. It comes with strong theoretical guarantees, and inspired approaches in deep Reinforcement Learning (RL). However, CPI itself has rarely been implemented, never with neural networks, and only experimented on toy problems. In this paper, we show how CPI can be practically combined with deep RL with discrete actions, in an off-policy manner. We also introduce adaptive mixture rates inspired by the theory. We experiment thoroughly the resulting algorithm on the simple Cartpole problem, and validate the proposed method on a representative subset of Atari games. Overall, this work suggests that revisiting classic ADP may lead to improved and more stable deep RL algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Learning to Constrain Policy Optimization with Virtual Trust RegionHung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen et al.NeurIPS 2022 · 5 citations
- Greedy Actor-Critic: A New Conditional Cross-Entropy Method for Policy ImprovementSamuel Neumann, Sungsu Lim, Ajin George Joseph, Yangchen Pan et al.ICLR 2023 · 3 citations
- MonoScale: Scaling Multi-Agent System with Monotonic ImprovementShuai Shao, Yixiang Liu, Bingwei Lu, Weinan ZhangICML 2026
Related papers
- Deep SPI: Safe Policy Improvement via World ModelsFlorent Delgrange, Raphaël Avalos, Willem RöpkeICLR 2026 · 4 citations
- Damped Anderson Mixing for Deep Reinforcement Learning: Acceleration, Convergence, and StabilizationKe Sun, Yafei Wang, Yi Liu, Yingnan Zhao et al.NeurIPS 2021 · 17 citations
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- Addressing Action Oscillations through Learning Policy InertiaChen Chen, Hongyao Tang, Jianye Hao, Wulong Liu et al.AAAI 2021 · 27 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
