Failure-Based Testing for Deep Reinforcement Learning Agents
Weibin Lin, Jiangtao Meng, Zheng Zheng
摘要
JIANGTAO MENG, Beihang University, China ZHENG ZHENG* , Beihang University, China Deep Reinforcement Learning (DRL) agents have been widely adopted across diverse domains to address challenging decision-making problems, such as autonomous driving and robotic control. Given that many of these applications are safety-and security-critical, rigorous testing of DRL agents is indispensable. Existing testing methods are typically guided by reward signals to detect failures. However, for well-trained agents, whose performance approaches optimal levels in standard operating conditions, reward signals remain generally high, making current methods ineffective at uncovering critical failures.
To address these challenges, we propose a novel failure-based method that leverages task-induced failure insights to enhance failure detection capability while reducing the number of tests required. Since DRL agents are inherently designed with human-defined tasks, they provide valuable cues about task difficulty. Intuitively, a DRL agent is more likely to fail when confronted with a more difficult task; therefore, PRT prioritizes these tasks. Building on this foundation, we propose Prior Random Testing, a black-box failure-based testing method that enables targeted prioritization while preserving the diversity of generated test cases. Guided by task-induced failure insights, PRT prioritizes failure-prone regions of the input domain, thereby facilitating efficient failure detection.
PRT is evaluated on four widely used benchmarks and compared with different state-of-the-art methods including fuzzing, search-based and generative-based methods. PRT ranks among the top performers in terms of both the cost of finding the first failure and the diversity of test cases. Notably, compared to random testing, PRT achieves better diversity and reduces the testing cost by over 50%.
CCS Concepts: • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 被引用 266 次
- MDPFuzz: testing models solving Markov decision processesQi Pang, Yuanyuan Yuan, Shuai WangISSTA 2022 · 被引用 37 次
- Generative Model-Based Testing on Decision-Making PoliciesZhuo Li, Xiongfei Wu, Derui Zhu, Mingfei Cheng 等ASE 2023 · 被引用 14 次
- Metamorphic relations via relaxations: an approach to obtain oracles for action-policy testingHasan Ferit Eniser, Timo P. Gros, Valentin Wüstholz, Jörg Hoffmann 等ISSTA 2022 · 被引用 14 次
- Curiosity-Driven Testing for Sequential Decision-Making ProcessJunda He, Zhou Yang, Jieke Shi, Chengran Yang 等ICSE 2024 · 被引用 6 次
相关 Paper
- Test-driven Reinforcement Learning in Continuous ControlZhao Yu, Xiuping Wu, Liangjun KeAAAI 2026
- : A Mutation Testing Pipeline for Deep Reinforcement Learning Based on Real FaultsDeepak-George Thomas, Matteo Biagiola, Nargiz Humbatova, Mohammad Wardat 等ICSE 2025 · 被引用 4 次
- Learning from the Test: Self-Referential Differential Testing for Deep RL AgentsJunda He, Jieke Shi, Zhou Yang, Mingfei Cheng 等ISSTA 2026
- GARL: Genetic Algorithm-Augmented Reinforcement Learning to Detect Violations in Marker-Based Autonomous Landing SystemsLinfeng Liang, Yao Deng, Kye Morton, Valtteri Kallinen 等ICSE 2025 · 被引用 9 次
- APIRL: Deep Reinforcement Learning for REST API FuzzingMyles Foley, Sergio MaffeisAAAI 2025 · 被引用 6 次
