Stabilizing PPO via Latent-Space Regularization and KDE-Driven Exploration
Meiyu Du, Yuqing Gao, Wei Wang
Abstract
Proximal Policy Optimization (PPO) is widely used in continuous-control tasks, yet its performance is often highly sensitive to training dynamics when neural networks approximate the policy and value functions. This paper introduces SPPO, a drop-in augmentation that preserves PPO's clipped objective and network topology while stabilizing actor-critic geometry via three mechanisms: (i) a Central Kernel Alignment (CKA)-based constraint on critic representations, (ii) a no-flip regularizer on actor updates, and (iii) Kernel Density Estimation (KDE)-driven advantage shaping. Theoretical analysis shows that these components tighten bounds on one-step bootstrapping error, improve expected directional alignment of action updates, and ensure nondecreasing occupancy mass over high-novelty regions. Experiments on standard continuouscontrol benchmarks demonstrate consistent gains over PPO and recent PPO stabilization methods. Ablation studies further quantify the contribution and complementary effects of each component. Additional training-dynamics analyses indicate that SPPO reduces instability and oscillations in both actor and critic updates, improving training stability and final performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f14da315-cee5-4b1f-ae9a-92fdae764c27Builds on23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 191 citations
Related papers
- Sampling Complexity of TD and PPO in RKHSLU ZOU, Wendi Ren, WEIZHONG ZHANG, Liang Ding et al.ICLR 2026 · 1 citation
- No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPOSkander Moalla, Andrea Miele, Daniil Pyatko, Razvan Pascanu et al.NeurIPS 2024 · 35 citations
- PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement LearningShunpeng Yang, Ben Liu, Hua ChenICLR 2026 · 6 citations
- Simple Policy OptimizationZhengpeng Xie, Qiang Zhang, Fan Yang, Marco Hutter et al.ICML 2025
- Off-Policy Proximal Policy OptimizationWenjia Meng, Qian Zheng, Gang Pan, Yilong YinAAAI 2023 · 27 citations
