Lune

ICML2026顶会

Stabilizing PPO via Latent-Space Regularization and KDE-Driven Exploration

Meiyu Du, Yuqing Gao, Wei Wang

出版方
2026年份

摘要

Proximal Policy Optimization (PPO) is widely used in continuous-control tasks, yet its performance is often highly sensitive to training dynamics when neural networks approximate the policy and value functions. This paper introduces SPPO, a drop-in augmentation that preserves PPO's clipped objective and network topology while stabilizing actor-critic geometry via three mechanisms: (i) a Central Kernel Alignment (CKA)-based constraint on critic representations, (ii) a no-flip regularizer on actor updates, and (iii) Kernel Density Estimation (KDE)-driven advantage shaping. Theoretical analysis shows that these components tighten bounds on one-step bootstrapping error, improve expected directional alignment of action updates, and ensure nondecreasing occupancy mass over high-novelty regions. Experiments on standard continuouscontrol benchmarks demonstrate consistent gains over PPO and recent PPO stabilization methods. Ablation studies further quantify the contribution and complementary effects of each component. Additional training-dynamics analyses indicate that SPPO reduces instability and oscillations in both actor and critic updates, improving training stability and final performance.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖