Metamorphic relations via relaxations: an approach to obtain oracles for action-policy testing
Hasan Ferit Eniser, Timo P. Gros, Valentin Wüstholz, Jörg Hoffmann, Maria Christakis
摘要
Testing is a promising way to gain trust in a learned action policy 𝜋, in particular if 𝜋 is a neural network. A "bug" in this context constitutes undesirable or fatal policy behavior, e.g., satisfying a failure condition. But how do we distinguish whether such behavior is due to bad policy decisions, or whether it is actually unavoidable under the given circumstances? This requires knowledge about optimal solutions, which defeats the scalability of testing. Related problems occur in software testing when the correct program output is not known. Metamorphic testing addresses this issue through metamorphic relations, specifying how a given change to the input should affect the output, thus providing an oracle for the correct output. Yet, how do we obtain such metamorphic relations for action policies? Here, we show that the well explored concept of relaxations in the Artificial Intelligence community can serve this purpose. In particular, if state 𝑠 ′ is a relaxation of state 𝑠, i.e., 𝑠 ′ is easier to solve than 𝑠, and 𝜋 fails on easier 𝑠 ′ but does not fail on harder 𝑠, then we know that 𝜋 contains a bug manifested on 𝑠 ′ . We contribute the first exploration of this idea in the context of failure testing of neural network policies 𝜋 learned by reinforcement learning in simulated environments. We design fuzzing strategies for test-case generation as well as metamorphic oracles leveraging simple, manually designed relaxations. In experiments on three single-agent games, our technology is able to effectively identify true bugs, i.e., avoidable failures of 𝜋, which has not been possible until now.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Curiosity-Driven Testing for Sequential Decision-Making ProcessJunda He, Zhou Yang, Jieke Shi, Chengran Yang 等ICSE 2024 · 被引用 6 次
- Policy Testing with MDPFuzz (Replicability Study)Quentin Mazouni, Helge Spieker, Arnaud Gotlieb, Mathieu AcherISSTA 2024 · 被引用 1 次
- Learning from the Test: Self-Referential Differential Testing for Deep RL AgentsJunda He, Jieke Shi, Zhou Yang, Mingfei Cheng 等ISSTA 2026
- Failure-Based Testing for Deep Reinforcement Learning AgentsWeibin Lin, Jiangtao Meng, Zheng ZhengFSE 2026
它引用的顶会 Paper3
- Structure-invariant testing for machine translationPinjia He, Clara Meister, Zhendong SuICSE 2020 · 被引用 84 次
- Importance-driven deep learning system testingSimos Gerasimou, Hasan Ferit Eniser, Alper Sen, Alper ÇakanICSE 2020 · 被引用 65 次
- Learning Generalized Relational Heuristic Networks for Model-Agnostic PlanningRushang Karia, Siddharth SrivastavaAAAI 2021 · 被引用 49 次
相关 Paper
- Stability-Aware Reinforcement Learning for Robust Class Integration Test Order GenerationYanru Ding, Yanmei Zhang, Guan Yuan, Shujuan Jiang 等AAAI 2026
- Natural Test Generation for Precise Testing of Question Answering SoftwareQingchao Shen, Junjie Chen, Jie M. Zhang, Haoyu Wang 等ASE 2022 · 被引用 27 次
- Abstraction-Aware Inference of Metamorphic RelationsAgustín Nolasco, Facundo Molina, Renzo Degiovanni, Alessandra Gorla 等FSE 2024 · 被引用 4 次
- Improving Deep Learning Framework Testing with Model-Level Metamorphic TestingYanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen 等ISSTA 2025 · 被引用 1 次
- MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling AnalysisCongying Xu, Hengcheng Zhu, Songqiang Chen, Jiarong Wu 等FSE 2026 · 被引用 1 次
