Perfect is the enemy of test oracle
Ali Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu, Reyhaneh Jabbarvand
Abstract
Automation of test oracles is one of the most challenging facets of software testing, but remains comparatively less addressed compared to automated test input generation. Test oracles rely on a ground-truth that can distinguish between the correct and buggy behavior to determine whether a test fails (detects a bug) or passes. What makes the oracle problem challenging and undecidable is the assumption that the ground-truth should know the exact expected, correct, or buggy behavior. However, we argue that one can still build an accurate oracle without knowing the exact correct or buggy behavior, but how these two might differ. This paper presents SEER, a learning-based approach that in the absence of test assertions or other types of oracle, can determine whether a unit test passes or fails on a given method under test (MUT). To build the groundtruth, SEER jointly embeds unit tests and the implementation of MUTs into a unified vector space, in such a way that the neural representation of tests are similar to that of MUTs they pass on them, but dissimilar to MUTs they fail on them. The classifier built on top of this vector representation serves as the oracle to generate "fail" labels, when test inputs detect a bug in MUT or "pass" labels, otherwise. Our extensive experiments on applying SEER to more than 5K unit tests from a diverse set of open-source Java projects show that the produced oracle is (1) effective in predicting the fail or pass labels, achieving an overall accuracy, precision, recall, and F1 measure of 93%, 86%, 94%, and 90%, (2) generalizable, predicting the labels for the unit test of projects that were not in training or validation set with negligible performance drop, and (3) efficient, detecting the existence of bugs in only 6.5 milliseconds on average. Moreover, by interpreting the neural model and looking at it beyond a closed-box solution, we confirm that the oracle is valid, i.e., it predicts the labels through learning relevant features.
• Software and its engineering → Software testing and debugging; • Computing methodologies → Neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2915598-2aa4-4087-aa86-b28e335b88f1Cited by top-tier papers5
- Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential PromptingTsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian et al.ASE 2023 · 51 citations
- Neural-Based Test Oracle Generation: A Large-Scale Evaluation and Lessons LearnedSoneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum et al.FSE 2023 · 30 citations
- AGORA: Automated Generation of Test Oracles for REST APIsJuan C. Alonso, Sergio Segura, Antonio Ruiz-CortésISSTA 2023 · 15 citations
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst et al.ASE 2025 · 3 citations
- SATORI: Static Test Oracle Generation for REST APIsJuan C. Alonso, Alberto Martin-Lopez, Sergio Segura, Gabriele Bavota et al.ASE 2025 · 2 citations
Builds on6
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota et al.ICSE 2020 · 96 citations
- Learning semantic program embeddings with graph interval neural networkYu Wang, Ke Wang, Fengjuan Gao, Linzhang WangOOPSLA 2020 · 61 citations
- Fuzzing Class SpecificationsFacundo Molina, Marcelo d'Amorim, Nazareno AguirreICSE 2022 · 25 citations
- Automated construction of energy test oracles for AndroidReyhaneh Jabbarvand, Forough Mehralian, Sam MalekFSE 2020 · 13 citations
Related papers
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 12 citations
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 92 citations
- Towards More Realistic Evaluation for Neural Test Oracle GenerationZhongxin Liu, Kui Liu, Xin Xia, Xiaohu YangISSTA 2023 · 24 citations
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney et al.ICSE 2023 · 45 citations
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program RepairHaoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu et al.ASE 2020 · 81 citations
