Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods
Jeong Woon Lee, Kyoleen Kwak, Daeho Kim, Hyoseok Hwang
摘要
Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physical deployment. Current approaches attempt to enforce smoothness by directly regularizing the policy's output. We argue that this approach treats the symptom rather than the cause. In this work, we theoretically establish that policy non-smoothness is fundamentally governed by the differential geometry of the critic. By applying implicit differentiation to the actor-critic objective, we prove that the sensitivity of the optimal policy is bounded by the ratio of the Q-function's mixed-partial derivative (noise sensitivity) to its action-space curvature (signal distinctness). To empirically validate this theoretical insight, we introduce PAVE (Policy-Aware Value-field Equalization), a critic-centric regularization framework that treats the critic as a scalar field and stabilizes its induced action-gradient field. PAVE rectifies the learning signal by minimizing the Q-gradient volatility while preserving local curvature. Experimental results demonstrate that PAVE achieves smoothness and robustness comparable to policy-side smoothness regularization methods, while maintaining competitive task performance, without modifying the actor.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Spectral Normalisation for Deep Reinforcement Learning: An Optimisation PerspectiveFlorin Gogianu, Tudor Berariu, Mihaela Rosca, Claudia Clopath 等ICML 2021 · 被引用 69 次
- Addressing Action Oscillations through Learning Policy InertiaChen Chen, Hongyao Tang, Jianye Hao, Wulong Liu 等AAAI 2021 · 被引用 27 次
- Towards Deeper Deep Reinforcement Learning with Spectral NormalizationJohan Bjorck, Carla P. Gomes, Kilian Q. WeinbergerNeurIPS 2021 · 被引用 26 次
- LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal ControlXujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li 等ICML 2023 · 被引用 19 次
- Enhancing Control Policy Smoothness by Aligning Actions with Predictions from Preceding StatesKyoleen Kwak, Hyoseok HwangAAAI 2026 · 被引用 1 次
相关 Paper
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional SmoothingFan Wu, Linyi Li, Zijian Huang, Yevgeniy Vorobeychik 等ICLR 2022 · 被引用 64 次
- Adaptive-Gradient Policy Optimization: Enhancing Policy Learning in Non-Smooth Differentiable SimulationsFeng Gao, Liangzhi Shi, Shenao Zhang, Zhaoran Wang 等ICML 2024 · 被引用 7 次
- Learning Value Functions in Deep Policy Gradients using Residual VarianceYannis Flet-Berliac, Reda Ouhamma, Odalric-Ambrym Maillard, Philippe PreuxICLR 2021 · 被引用 16 次
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 被引用 10 次
- Stabilizing PPO via Latent-Space Regularization and KDE-Driven ExplorationMeiyu Du, Yuqing Gao, Wei WangICML 2026
