Learning from the Test: Self-Referential Differential Testing for Deep RL Agents
Junda He, Jieke Shi, Zhou Yang, Mingfei Cheng, David Lo
Abstract
Deep Reinforcement Learning (DRL) has achieved significant success in complex decision-making problems. As DRL systems are increasingly deployed in real-world applications, ensuring their quality and reliability is paramount. Current works primarily focus on detecting safety-critical failures, often neglecting policy optimality, which can lead to reduced efficiency, user distrust, and economic losses. This oversight, compounded by the inherent “testing oracle problem” for optimality, leaves a significant gap in comprehensively evaluating DRL systems. To address this gap, we propose Delta (Differential Testing for DRL Agents), a novel and comprehensive framework that automatically identifies both safety-critical and optimality bugs in DRL agents. Delta employs a two-phase approach: (1) Safety Testing, where the Agent Under Test (AUT) is evaluated for catastrophic failures while collecting data from its decision-making policy, and (2) Optimality Testing, where this collected data from the prior phase is used to train a challenger agent via Offline Reinforcement Learning. Differential testing is then performed by comparing the challenger agent against the AUT; instances where the challenger achieves higher cumulative rewards indicate optimality issues in the AUT. We demonstrate Delta’s effectiveness across five environments. We investigate the effectiveness of three offline RL algorithms (BC, BCQ, and CQL) in generating challenger agents. Experimental results demonstrate that safety testing datasets are valuable for training competent DRL agents. Challenger agents trained with BCQ proved most effective for identifying optimality issues within the framework of Delta. Across the five environments, Delta uncovered an average of 2,518 optimality issues, outperforming the baseline methods by 50.2%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b068b0d-796b-4713-aee2-033882451ff7Builds on23
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong et al.NeurIPS 2021 · 207 citations
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 150 citations
- Making Efficient Use of Demonstrations to Solve Hard Exploration ProblemsÇaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil et al.ICLR 2020 · 97 citations
Related papers
- A Test Oracle for Reinforcement Learning Software Based on Lyapunov Stability Control TheoryShiyu Zhang, Haoyang Song, Qixin Wang, Henghua Shen et al.ICSE 2025 · 1 citation
- Many-Objective Reinforcement Learning for Online Testing of DNN-Enabled SystemsFitash Ul Haq, Donghwan Shin, Lionel C. BriandICSE 2023 · 35 citations
- Test Where Decisions Matter: Importance-driven Testing for Deep Reinforcement LearningStefan Pranger, Hana Chockler, Martin Tappler, Bettina KönighoferNeurIPS 2024 · 7 citations
- Real-DRL: Teach and Learn at RuntimeYanbing Mao, Yihao Cai, Lui ShaNeurIPS 2025 · 2 citations
- Deeply Reinforcing Android GUI Testing with Deep Reinforcement LearningYuanhong Lan, Yifei Lu, Zhong Li, Minxue Pan et al.ICSE 2024 · 21 citations
