Metamorphic relations via relaxations: an approach to obtain oracles for action-policy testing
Hasan Ferit Eniser, Timo P. Gros, Valentin Wüstholz, Jörg Hoffmann, Maria Christakis
Abstract
Testing is a promising way to gain trust in a learned action policy 𝜋, in particular if 𝜋 is a neural network. A "bug" in this context constitutes undesirable or fatal policy behavior, e.g., satisfying a failure condition. But how do we distinguish whether such behavior is due to bad policy decisions, or whether it is actually unavoidable under the given circumstances? This requires knowledge about optimal solutions, which defeats the scalability of testing. Related problems occur in software testing when the correct program output is not known. Metamorphic testing addresses this issue through metamorphic relations, specifying how a given change to the input should affect the output, thus providing an oracle for the correct output. Yet, how do we obtain such metamorphic relations for action policies? Here, we show that the well explored concept of relaxations in the Artificial Intelligence community can serve this purpose. In particular, if state 𝑠 ′ is a relaxation of state 𝑠, i.e., 𝑠 ′ is easier to solve than 𝑠, and 𝜋 fails on easier 𝑠 ′ but does not fail on harder 𝑠, then we know that 𝜋 contains a bug manifested on 𝑠 ′ . We contribute the first exploration of this idea in the context of failure testing of neural network policies 𝜋 learned by reinforcement learning in simulated environments. We design fuzzing strategies for test-case generation as well as metamorphic oracles leveraging simple, manually designed relaxations. In experiments on three single-agent games, our technology is able to effectively identify true bugs, i.e., avoidable failures of 𝜋, which has not been possible until now.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 769cdbcf-d17e-47c0-8da0-db918d6a95c5Cited by top-tier papers4
- Curiosity-Driven Testing for Sequential Decision-Making ProcessJunda He, Zhou Yang, Jieke Shi, Chengran Yang et al.ICSE 2024 · 6 citations
- Policy Testing with MDPFuzz (Replicability Study)Quentin Mazouni, Helge Spieker, Arnaud Gotlieb, Mathieu AcherISSTA 2024 · 1 citation
- Learning from the Test: Self-Referential Differential Testing for Deep RL AgentsJunda He, Jieke Shi, Zhou Yang, Mingfei Cheng et al.ISSTA 2026
- Failure-Based Testing for Deep Reinforcement Learning AgentsWeibin Lin, Jiangtao Meng, Zheng ZhengFSE 2026
Builds on3
- Structure-invariant testing for machine translationPinjia He, Clara Meister, Zhendong SuICSE 2020 · 84 citations
- Importance-driven deep learning system testingSimos Gerasimou, Hasan Ferit Eniser, Alper Sen, Alper ÇakanICSE 2020 · 65 citations
- Learning Generalized Relational Heuristic Networks for Model-Agnostic PlanningRushang Karia, Siddharth SrivastavaAAAI 2021 · 49 citations
Related papers
- Stability-Aware Reinforcement Learning for Robust Class Integration Test Order GenerationYanru Ding, Yanmei Zhang, Guan Yuan, Shujuan Jiang et al.AAAI 2026
- Natural Test Generation for Precise Testing of Question Answering SoftwareQingchao Shen, Junjie Chen, Jie M. Zhang, Haoyu Wang et al.ASE 2022 · 27 citations
- Abstraction-Aware Inference of Metamorphic RelationsAgustín Nolasco, Facundo Molina, Renzo Degiovanni, Alessandra Gorla et al.FSE 2024 · 4 citations
- Improving Deep Learning Framework Testing with Model-Level Metamorphic TestingYanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen et al.ISSTA 2025 · 1 citation
- MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling AnalysisCongying Xu, Hengcheng Zhu, Songqiang Chen, Jiarong Wu et al.FSE 2026 · 1 citation
