A Test Oracle for Reinforcement Learning Software Based on Lyapunov Stability Control Theory
Shiyu Zhang, Haoyang Song, Qixin Wang, Henghua Shen, Yu Pei
Abstract
Reinforcement Learning (RL) has gained significant attention in recent years. As RL software becomes more complex and infiltrates critical application domains, ensuring its quality and correctness becomes increasingly important. An indispensable aspect of software quality/correctness insurance is testing. However, testing RL software faces unique challenges compared to testing traditional software, due to the difficulty on defining the outputs' correctness. This leads to the RL test oracle problem. Current approaches to testing RL software often rely on human oracles, i.e. convening human experts to judge the correctness of RL software outputs. This heavily depends on the availability and quality (including the experiences, subjective states, etc.) of the human experts, and cannot be fully automated. In this paper, we propose a novel approach to design test oracles for RL software by leveraging the Lyapunov stability control theory. By incorporating Lyapunov stability concepts to guide RL training, we hypothesize that a correctly implemented RL software shall output an agent that respects Lyapunov stability control theories. Based on this heuristics, we propose a Lyapunov stability control theory based oracle, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>, for testing RL software. We conduct extensive experiments over representative RL algorithms and RL software bugs to evaluate our proposed oracle. The results show that our proposed oracle can outperform the human oracle in most metrics. Particularly, LPEA <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> outperforms the human oracle by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>, and 31.7 % respectively on accuracy, precision, recall, F1 score, true positive rate, true negative rate, false positive rate, false negative rate, and ROC curve's AUC; and LPEA (<tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>) outperforms the human oracle by <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>, <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex>, and 26.0 % respectively on these metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0cef36d2-4c49-4589-aec6-47c6bb309ef0Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Learning from the Test: Self-Referential Differential Testing for Deep RL AgentsJunda He, Jieke Shi, Zhou Yang, Mingfei Cheng et al.ISSTA 2026
- Metamorphic relations via relaxations: an approach to obtain oracles for action-policy testingHasan Ferit Eniser, Timo P. Gros, Valentin Wüstholz, Jörg Hoffmann et al.ISSTA 2022 · 14 citations
- : A Mutation Testing Pipeline for Deep Reinforcement Learning Based on Real FaultsDeepak-George Thomas, Matteo Biagiola, Nargiz Humbatova, Mohammad Wardat et al.ICSE 2025 · 4 citations
- Perfect is the enemy of test oracleAli Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu, Reyhaneh JabbarvandFSE 2022 · 23 citations
- Reinforcement learning based curiosity-driven testing of Android applicationsMinxue Pan, An Huang, Guoxin Wang, Tian Zhang et al.ISSTA 2020 · 166 citations
