Certifying Stability of Reinforcement Learning Policies using Generalized Lyapunov Functions
Kehan Long, Jorge Cortés, Nikolay Atanasov
摘要
Establishing stability certificates for closed-loop systems under reinforcement learning (RL) policies is essential to move beyond empirical performance and offer guarantees of system behavior. Classical Lyapunov methods require a strict stepwise decrease in the Lyapunov function but such certificates are difficult to construct for learned policies. The RL value function is a natural candidate but it is not well understood how it can be adapted for this purpose. To gain intuition, we first study the linear quadratic regulator (LQR) problem and make two key observations. First, a Lyapunov function can be obtained from the value function of an LQR policy by augmenting it with a residual term related to the system dynamics and stage cost. Second, the classical Lyapunov decrease requirement can be relaxed to a generalized Lyapunov condition requiring only decrease on average over multiple time steps. Using this intuition, we consider the nonlinear setting and formulate an approach to learn generalized Lyapunov functions by augmenting RL value functions with neural network residual terms. Our approach successfully certifies the stability of RL policies trained on Gymnasium and DeepMind Control benchmarks. We also extend our method to jointly train neural controllers and stability certificates using a multi-step Lyapunov loss, resulting in larger certified inner approximations of the region of attraction compared to the classical Lyapunov approach. Overall, our formulation enables stability certification for a broad class of systems with learned policies by making certificates easier to construct, thereby bridging classical control theory and modern learning-based methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Stability Verification in Stochastic Control Systems via Neural Network SupermartingalesMathias Lechner, Dorde Zikelic, Krishnendu Chatterjee, Thomas A. HenzingerAAAI 2022 · 被引用 45 次
- Neural Lyapunov Control for Discrete-Time SystemsJunlin Wu, Andrew Clark, Yiannis Kantaros, Yevgeniy VorobeychikNeurIPS 2023 · 被引用 61 次
- Neural Lyapunov Control of Unknown Nonlinear Systems with Stability GuaranteesRuikun Zhou, Thanin Quartz, Hans De Sterck, Jun LiuNeurIPS 2022 · 被引用 109 次
- Neural Control and Certificate Repair via Runtime MonitoringEmily Yu, Dorde Zikelic, Thomas A. HenzingerAAAI 2025 · 被引用 5 次
- Two‑Stage Learning of Stabilizing Neural Controllers via Zubov Sampling and Iterative Domain ExpansionHaoyu Li, Xiangru Zhong, Bin Hu, Huan ZhangNeurIPS 2025 · 被引用 11 次
