Robustness Verification of Deep Reinforcement Learning Based Control Systems Using Reward Martingales
Dapeng Zhi, Peixin Wang, Cheng Chen, Min Zhang
摘要
Deep Reinforcement Learning (DRL) has gained prominence as an effective approach for control systems. However, its practical deployment is impeded by state perturbations that can severely impact system performance. Addressing this critical challenge requires robustness verification about system performance, which involves tackling two quantitative questions: (i) how to establish guaranteed bounds for expected cumulative rewards, and (ii) how to determine tail bounds for cumulative rewards. In this work, we present the first approach for robustness verification of DRL-based control systems by introducing reward martingales, which offer a rigorous mathematical foundation to characterize the impact of state perturbations on system performance in terms of cumulative rewards. Our verified results provide provably quantitative certificates for the two questions. We then show that reward martingales can be implemented and trained via neural networks, against different types of control policies. Experimental results demonstrate that our certified bounds tightly enclose simulation outcomes on various DRL-based control systems, indicating the effectiveness and generality of the proposed approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Toward A Thousand Lights: Decentralized Deep Reinforcement Learning for Large-Scale Traffic Signal ControlChacha Chen, Hua Wei, Nan Xu, Guanjie Zheng 等AAAI 2020 · 被引用 450 次
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang 等NeurIPS 2020 · 被引用 415 次
- Robust Deep Reinforcement Learning through Adversarial LossTuomas P. Oikarinen, Wang Zhang, Alexandre Megretski, Luca Daniel 等NeurIPS 2021 · 被引用 134 次
- Policy Smoothing for Provably Robust Reinforcement LearningAounon Kumar, Alexander Levine, Soheil FeiziICLR 2022 · 被引用 62 次
相关 Paper
- Stability Verification in Stochastic Control Systems via Neural Network SupermartingalesMathias Lechner, Dorde Zikelic, Krishnendu Chatterjee, Thomas A. HenzingerAAAI 2022 · 被引用 45 次
- Unifying Qualitative and Quantitative Safety Verification of DNN-Controlled SystemsDapeng Zhi, Peixin Wang, Si Liu, C.-H. Luke Ong 等CAV 2024 · 被引用 11 次
- Neural Network Control Policy Verification With Persistent Adversarial PerturbationYuh-Shyang Wang, Lily Weng, Luca DanielICML 2020 · 被引用 7 次
- Improve Robustness of Reinforcement Learning against Observation Perturbations via l∞ Lipschitz Policy NetworksBuqing Nie, Jingtian Ji, Yangqing Fu, Yue GaoAAAI 2024 · 被引用 10 次
- Trainify: A CEGAR-Driven Training and Verification Framework for Safe Deep Reinforcement LearningPeng Jin, Jiaxu Tian, Dapeng Zhi, Xuejun Wen 等CAV 2022 · 被引用 28 次
