LipsNet: A Smooth and Robust Neural Network with Adaptive Lipschitz Constant for High Accuracy Optimal Control
Xujie Song, Jingliang Duan, Wenxuan Wang, Shengbo Eben Li, Chen Chen, Bo Cheng, Bo Zhang, Junqing Wei, Xiaoming Simon Wang
摘要
Deep reinforcement learning (RL) is a powerful approach for solving optimal control problems. However, RL-trained policies often suffer from the action fluctuation problem, where the consecutive actions significantly differ despite only slight state variations. This problem results in mechanical components' wear and tear and poses safety hazards. The action fluctuation is caused by the high Lipschitz constant of actor networks. To address this problem, we propose a neural network named LipsNet. We propose the Multi-dimensional Gradient Normalization (MGN) method, to constrain the Lipschitz constant of networks with multi-dimensional input and output. Benefiting from MGN, LipsNet achieves Lipschitz continuity, allowing smooth actions while preserving control performance by adjusting Lipschitz constant. LipsNet addresses the action fluctuation problem at network level rather than algorithm level, which can serve as actor networks in most RL algorithms, making it more flexible and user-friendly than previous works. Experiments demonstrate that LipsNet has good landscape smoothness and noise robustness, resulting in significantly smoother action compared to the Multilayer Perceptron.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Enhancing Control Policy Smoothness by Aligning Actions with Predictions from Preceding StatesKyoleen Kwak, Hyoseok HwangAAAI 2026 · 被引用 1 次
- Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic MethodsJeong Woon Lee, Kyoleen Kwak, Daeho Kim, Hyoseok HwangICML 2026 · 被引用 1 次
- Knowledgeable Language Models as Black-Box Optimizers for Personalized MedicineMichael S. Yao, Osbert Bastani, Alma Andersson, Tommaso Biancalani 等ICLR 2026
- ODE-based Smoothing Neural Network for Reinforcement Learning TasksYinuo Wang, Wenxuan Wang, Xujie Song, Tong Liu 等ICLR 2025
- LipsNet++: Unifying Filter and Controller into a Policy NetworkXujie Song, Liangfa Chen, Tong Liu, Wenxuan Wang 等ICML 2025
它引用的顶会 Paper9
- Deep Reinforcement Learning with Robust and Smooth PolicyQianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang 等ICML 2020 · 被引用 95 次
- Reachability Constrained Reinforcement LearningDongjie Yu, Haitong Ma, Sheng-bo Li, Jianyu ChenICML 2022 · 被引用 90 次
- Gradient Normalization for Generative Adversarial NetworksYi-Lun Wu, Hong-Han Shuai, Zhi Rui Tam, Hong-Yu ChiuICCV 2021 · 被引用 78 次
- Spectral Normalisation for Deep Reinforcement Learning: An Optimisation PerspectiveFlorin Gogianu, Tudor Berariu, Mihaela Rosca, Claudia Clopath 等ICML 2021 · 被引用 69 次
- TAAC: Temporally Abstract Actor-Critic for Continuous ControlHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2021 · 被引用 30 次
相关 Paper
- Fractal Landscapes in Policy OptimizationTao Wang, Sylvia L. Herbert, Sicun GaoNeurIPS 2023 · 被引用 10 次
- Action Manifold Smoothing: A Lipschitz Pathway Perspective on High-Dimensional Reinforcement LearningZhihao LinICML 2026
- Leveraging Constraint Violation Signals for Action Constrained Reinforcement LearningJanaka Chathuranga Brahmanage, Jiajing Ling, Akshat KumarAAAI 2025 · 被引用 2 次
- Improve Robustness of Reinforcement Learning against Observation Perturbations via l∞ Lipschitz Policy NetworksBuqing Nie, Jingtian Ji, Yangqing Fu, Yue GaoAAAI 2024 · 被引用 10 次
- Model-Based Reparameterization Policy Gradient Methods: Theory and Practical AlgorithmsShenao Zhang, Boyi Liu, Zhaoran Wang, Tuo ZhaoNeurIPS 2023 · 被引用 8 次
