Boosting Verification of Deep Reinforcement Learning via Piece-Wise Linear Decision Neural Networks
Jiaxu Tian, Dapeng Zhi, Si Liu, Peixin Wang, Cheng Chen, Min Zhang
摘要
Formally verifying deep reinforcement learning (DRL) systems su ff ers from both inaccurate verification results and limited scalability. The major obstacle lies in the large overestimation introduced inherently during training and then transforming the inexplicable decision-making models i.e., deep neural networks (DNNs), into easy-to-verify models. In this paper, we propose an inverse transform-then-train approach, which first encodes a DNN into an equivalent set of e ffi ciently and tightly verifiable linear control policies and then optimizes them via reinforcement learning. We accompany our inverse approach with a novel neural network model called piece-wise linear decision neural networks (PLDNNs), which are compatible with most existing DRL training algorithms with comparable performance against conventional DNNs. Our extensive experiments show that, compared to DNN-based DRL systems, PLDNN-based systems can be more e ffi ciently and tightly verified with up to 438 times speedup and a significant reduction in overestimation. In particular, even a complex 12-dimensional DRL system is e ffi ciently verified with up to 7 times deeper computation steps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Verification of Deep Convolutional Neural Networks Using ImageStarsHoang-Dung Tran, Stanley Bak, Weiming Xiang, Taylor T. JohnsonCAV 2020 · 被引用 122 次
- Learning to Synthesize Programs as Interpretable and Generalizable PoliciesDweep Trivedi, Jesse Zhang, Shao-Hua Sun, Joseph J. LimNeurIPS 2021 · 被引用 104 次
- Verification of Neural-Network Control Systems by Integrating Taylor Models and ZonotopesChristian Schilling, Marcelo Forets, Sebastián GuadalupeAAAI 2022 · 被引用 48 次
- Trainify: A CEGAR-Driven Training and Verification Framework for Safe Deep Reinforcement LearningPeng Jin, Jiaxu Tian, Dapeng Zhi, Xuejun Wen 等CAV 2022 · 被引用 28 次
- Interval universal approximation for neural networksZi Wang, Aws Albarghouthi, Gautam Prakriya, Somesh JhaPOPL 2022 · 被引用 18 次
相关 Paper
- Safe DNN-type Controller Synthesis for Nonlinear Systems via Meta Reinforcement LearningHanrui Zhao, Xia Zeng, Niuniu Qi, Zhengfeng Yang 等DAC 2023 · 被引用 4 次
- Safe Controller Synthesis for Nonlinear Systems via Reinforcement Learning and PAC ApproximationXia Zeng, Banglong Liu, Zhenbing Zeng, Zhiming Liu 等DAC 2024 · 被引用 1 次
- An Iterative Scheme of Safe Reinforcement Learning for Nonlinear Systems via Barrier Certificate GenerationZhengfeng Yang, Yidan Zhang, Wang Lin, Xia Zeng 等CAV 2021 · 被引用 15 次
- Unifying Qualitative and Quantitative Safety Verification of DNN-Controlled SystemsDapeng Zhi, Peixin Wang, Si Liu, C.-H. Luke Ong 等CAV 2024 · 被引用 11 次
- Verifying learning-augmented systemsTomer Eliyahu, Yafim Kazak, Guy Katz, Michael SchapiraSIGCOMM 2021 · 被引用 45 次
