Failure-Based Testing for Deep Reinforcement Learning Agents
Weibin Lin, Jiangtao Meng, Zheng Zheng
Abstract
JIANGTAO MENG, Beihang University, China ZHENG ZHENG* , Beihang University, China Deep Reinforcement Learning (DRL) agents have been widely adopted across diverse domains to address challenging decision-making problems, such as autonomous driving and robotic control. Given that many of these applications are safety-and security-critical, rigorous testing of DRL agents is indispensable. Existing testing methods are typically guided by reward signals to detect failures. However, for well-trained agents, whose performance approaches optimal levels in standard operating conditions, reward signals remain generally high, making current methods ineffective at uncovering critical failures.
To address these challenges, we propose a novel failure-based method that leverages task-induced failure insights to enhance failure detection capability while reducing the number of tests required. Since DRL agents are inherently designed with human-defined tasks, they provide valuable cues about task difficulty. Intuitively, a DRL agent is more likely to fail when confronted with a more difficult task; therefore, PRT prioritizes these tasks. Building on this foundation, we propose Prior Random Testing, a black-box failure-based testing method that enables targeted prioritization while preserving the diversity of generated test cases. Guided by task-induced failure insights, PRT prioritizes failure-prone regions of the input domain, thereby facilitating efficient failure detection.
PRT is evaluated on four widely used benchmarks and compared with different state-of-the-art methods including fuzzing, search-based and generative-based methods. PRT ranks among the top performers in terms of both the cost of finding the first failure and the diversity of test cases. Notably, compared to random testing, PRT achieves better diversity and reduces the testing cost by over 50%.
CCS Concepts: • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad1192d2-8ce3-4b01-be99-aa25a6e3a74eBuilds on7
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- MDPFuzz: testing models solving Markov decision processesQi Pang, Yuanyuan Yuan, Shuai WangISSTA 2022 · 37 citations
- Generative Model-Based Testing on Decision-Making PoliciesZhuo Li, Xiongfei Wu, Derui Zhu, Mingfei Cheng et al.ASE 2023 · 14 citations
- Metamorphic relations via relaxations: an approach to obtain oracles for action-policy testingHasan Ferit Eniser, Timo P. Gros, Valentin Wüstholz, Jörg Hoffmann et al.ISSTA 2022 · 14 citations
- Curiosity-Driven Testing for Sequential Decision-Making ProcessJunda He, Zhou Yang, Jieke Shi, Chengran Yang et al.ICSE 2024 · 6 citations
Related papers
- Test-driven Reinforcement Learning in Continuous ControlZhao Yu, Xiuping Wu, Liangjun KeAAAI 2026
- : A Mutation Testing Pipeline for Deep Reinforcement Learning Based on Real FaultsDeepak-George Thomas, Matteo Biagiola, Nargiz Humbatova, Mohammad Wardat et al.ICSE 2025 · 4 citations
- Learning from the Test: Self-Referential Differential Testing for Deep RL AgentsJunda He, Jieke Shi, Zhou Yang, Mingfei Cheng et al.ISSTA 2026
- GARL: Genetic Algorithm-Augmented Reinforcement Learning to Detect Violations in Marker-Based Autonomous Landing SystemsLinfeng Liang, Yao Deng, Kye Morton, Valtteri Kallinen et al.ICSE 2025 · 9 citations
- APIRL: Deep Reinforcement Learning for REST API FuzzingMyles Foley, Sergio MaffeisAAAI 2025 · 6 citations
