Learning from the Test: Self-Referential Differential Testing for Deep RL Agents
Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng, David Lo
摘要
Deep Reinforcement Learning (DRL) has achieved significant success in complex decision-making problems. As DRL systems are increasingly deployed in real-world applications, ensuring their quality and reliability is paramount. Current works primarily focus on detecting safety-critical failures, often neglecting policy optimality, which can lead to reduced efficiency, user distrust, and economic losses. This oversight, compounded by the inherent “testing oracle problem” for optimality, leaves a significant gap in comprehensively evaluating DRL systems. To address this gap, we propose Delta (Differential Testing for DRL Agents), a novel and comprehensive framework that automatically identifies both safety-critical and optimality bugs in DRL agents. Delta employs a two-phase approach: (1) Safety Testing, where the Agent Under Test (AUT) is evaluated for catastrophic failures while collecting data from its decision-making policy, and (2) Optimality Testing, where this collected data from the prior phase is used to train a challenger agent via Offline Reinforcement Learning. Differential testing is then performed by comparing the challenger agent against the AUT; instances where the challenger achieves higher cumulative rewards indicate optimality issues in the AUT. We demonstrate Delta’s effectiveness across five environments. We investigate the effectiveness of three offline RL algorithms (BC, BCQ, and CQL) in generating challenger agents. Experimental results demonstrate that safety testing datasets are valuable for training competent DRL agents. Challenger agents trained with BCQ proved most effective for identifying optimality issues within the framework of Delta. Across the five environments, Delta uncovered an average of 2,518 optimality issues, outperforming the baseline methods by 50.2%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong 等NeurIPS 2021 · 被引用 207 次
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 被引用 150 次
- Making Efficient Use of Demonstrations to Solve Hard Exploration ProblemsÇaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil 等ICLR 2020 · 被引用 97 次
相关 Paper
- A Test Oracle for Reinforcement Learning Software Based on Lyapunov Stability Control TheoryShiyu Zhang, Haoyang Song, Qixin Wang, Henghua Shen 等ICSE 2025 · 被引用 1 次
- Many-Objective Reinforcement Learning for Online Testing of DNN-Enabled SystemsFitash Ul Haq, Donghwan Shin, Lionel C. BriandICSE 2023 · 被引用 35 次
- Test Where Decisions Matter: Importance-driven Testing for Deep Reinforcement LearningStefan Pranger, Hana Chockler, Martin Tappler, Bettina KönighoferNeurIPS 2024 · 被引用 7 次
- Real-DRL: Teach and Learn at RuntimeYanbing Mao, Yihao Cai, Lui ShaNeurIPS 2025 · 被引用 2 次
- Deeply Reinforcing Android GUI Testing with Deep Reinforcement LearningYuanhong Lan, Yifei Lu, Zhong Li, Minxue Pan 等ICSE 2024 · 被引用 21 次
