Lune

ICML2026Top-tier venue

Stabilizing PPO via Latent-Space Regularization and KDE-Driven Exploration

Meiyu Du, Yuqing Gao, Wei Wang

2026Year

Abstract

Proximal Policy Optimization (PPO) is widely used in continuous-control tasks, yet its performance is often highly sensitive to training dynamics when neural networks approximate the policy and value functions. This paper introduces SPPO, a drop-in augmentation that preserves PPO's clipped objective and network topology while stabilizing actor-critic geometry via three mechanisms: (i) a Central Kernel Alignment (CKA)-based constraint on critic representations, (ii) a no-flip regularizer on actor updates, and (iii) Kernel Density Estimation (KDE)-driven advantage shaping. Theoretical analysis shows that these components tighten bounds on one-step bootstrapping error, improve expected directional alignment of action updates, and ensure nondecreasing occupancy mass over high-novelty regions. Experiments on standard continuouscontrol benchmarks demonstrate consistent gains over PPO and recent PPO stabilization methods. Ablation studies further quantify the contribution and complementary effects of each component. Additional training-dynamics analyses indicate that SPPO reduces instability and oscillations in both actor and critic updates, improving training stability and final performance.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f14da315-cee5-4b1f-ae9a-92fdae764c27

Builds on23

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines