Harmonized Dual Policy Improvement for Modelic Reinforcement Learning
Guojian Zhan, Likun Wang, Feihong Zhang, Yang Guan, Shengbo Li
Abstract
Policy-planner bootstrapping has emerged as a powerful paradigm in model-based reinforcement learning (MBRL). We formalize this process as a dual policy improvement mechanism synergizing: (i) exploitative improvement via off-policy Q-maximization, and (ii) lookahead improvement via planner alignment. While we theoretically prove that these improvements anchor to the same optimum, the practical training process inevitably encounters gradient inconsistency. Exacerbated by approximation inaccuracies and non-stationary data, this inconsistency induces destructive interference in policy updates, destabilizing the bootstrapping loop and leading to suboptimal convergence. To address this, we propose harmonized dual policy improvement (HDPI), a gradient-level framework that reconciles exploitative and lookahead improvements through a harmonized optimization scheme. This scheme effectively maximizes the worst-case inner product between the harmonized update and the original gradients, ensuring directional consistency and stabilizing policy evolution. Extensive empirical evaluations on 14 challenging tasks from the DeepMind Control Suite and the Humanoid-Bench demonstrate that HDPI significantly enhances training stability and asymptotic performance, outperforming a wide range of strong baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
Related papers
- Bootstrap Off-policy with World ModelGuojian Zhan, Likun Wang, Xiangteng Zhang, Jiaxin Gao et al.NeurIPS 2025 · 9 citations
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 18 citations
- Gradient Information Matters in Policy Optimization by Back-propagating through ModelChongchong Li, Yue Wang, Wei Chen, Yuting Liu et al.ICLR 2022 · 10 citations
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution CorrectionAviral Kumar, Abhishek Gupta, Sergey LevineNeurIPS 2020 · 124 citations
- A KL-regularization framework for learning to plan with adaptive priorsÁlvaro Serra-Gómez, Daniel Jarne Ornia, Dhruva Tirumala, Thomas M MoerlandICML 2026
