Harmonized Dual Policy Improvement for Modelic Reinforcement Learning
Guojian Zhan, Likun Wang, Feihong Zhang, Yang Guan, Shengbo Li
摘要
Policy-planner bootstrapping has emerged as a powerful paradigm in model-based reinforcement learning (MBRL). We formalize this process as a dual policy improvement mechanism synergizing: (i) exploitative improvement via off-policy Q-maximization, and (ii) lookahead improvement via planner alignment. While we theoretically prove that these improvements anchor to the same optimum, the practical training process inevitably encounters gradient inconsistency. Exacerbated by approximation inaccuracies and non-stationary data, this inconsistency induces destructive interference in policy updates, destabilizing the bootstrapping loop and leading to suboptimal convergence. To address this, we propose harmonized dual policy improvement (HDPI), a gradient-level framework that reconciles exploitative and lookahead improvements through a harmonized optimization scheme. This scheme effectively maximizes the worst-case inner product between the harmonized update and the original gradients, ensuring directional consistency and stabilizing policy evolution. Extensive empirical evaluations on 14 challenging tasks from the DeepMind Control Suite and the Humanoid-Bench demonstrate that HDPI significantly enhances training stability and asymptotic performance, outperforming a wide range of strong baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
相关 Paper
- Bootstrap Off-policy with World ModelGuojian Zhan, Likun Wang, Xiangteng Zhang, Jiaxin Gao 等NeurIPS 2025 · 被引用 9 次
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
- Gradient Information Matters in Policy Optimization by Back-propagating through ModelChongchong Li, Yue Wang, Wei Chen, Yuting Liu 等ICLR 2022 · 被引用 10 次
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution CorrectionAviral Kumar, Abhishek Gupta, Sergey LevineNeurIPS 2020 · 被引用 124 次
- A KL-regularization framework for learning to plan with adaptive priorsÁlvaro Serra-Gómez, Daniel Jarne Ornia, Dhruva Tirumala, Thomas M MoerlandICML 2026
