Policy Search by Target Distribution Learning for Continuous Control
Chuheng Zhang, Yuanqi Li, Jian Li
摘要
It is known that existing policy gradient methods (such as vanilla policy gradient, PPO, A2C) may suffer from overly large gradients when the current policy is close to deterministic, leading to an unstable training process. We show that such instability can happen even in a very simple environment. To address this issue, we propose a new method, called target distribution learning (TDL), for policy improvement in reinforcement learning. TDL alternates between proposing a target distribution and training the policy network to approach the target distribution. TDL is more effective in constraining the KL divergence between updated policies, and hence leads to more stable policy improvements over iterations. Our experiments show that TDL algorithms perform comparably to (or better than) state-of-the-art algorithms for most continuous control tasks in the MuJoCo environment while being more stable in training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Stabilizing Policy Gradient Methods via Reward ProfilingShihab Ahmed, El Houcine Bergou, Yue Wang, Aritra DuttaAAAI 2026
- Improving Value Estimation Critically Enhances Vanilla Policy GradientTao Wang, Ruipeng Zhang, Sicun GaoICML 2025
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta 等ICML 2020 · 被引用 49 次
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
- PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation ModelsJeongjae Lee, Jong Chul YeICLR 2026 · 被引用 2 次
