Policy Search by Target Distribution Learning for Continuous Control
Chuheng Zhang, Yuanqi Li, Jian Li
Abstract
It is known that existing policy gradient methods (such as vanilla policy gradient, PPO, A2C) may suffer from overly large gradients when the current policy is close to deterministic, leading to an unstable training process. We show that such instability can happen even in a very simple environment. To address this issue, we propose a new method, called target distribution learning (TDL), for policy improvement in reinforcement learning. TDL alternates between proposing a target distribution and training the policy network to approach the target distribution. TDL is more effective in constraining the KL divergence between updated policies, and hence leads to more stable policy improvements over iterations. Our experiments show that TDL algorithms perform comparably to (or better than) state-of-the-art algorithms for most continuous control tasks in the MuJoCo environment while being more stable in training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e87b8e3-a141-492c-9c4d-6be7d9487b01Related papers
- Stabilizing Policy Gradient Methods via Reward ProfilingShihab Ahmed, El Houcine Bergou, Yue Wang, Aritra DuttaAAAI 2026
- Improving Value Estimation Critically Enhances Vanilla Policy GradientTao Wang, Ruipeng Zhang, Sicun GaoICML 2025
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta et al.ICML 2020 · 49 citations
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
- PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation ModelsJeongjae Lee, Jong Chul YeICLR 2026 · 2 citations
