Perfect is the enemy of test oracle
Ali Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu, Reyhaneh Jabbarvand
摘要
Automation of test oracles is one of the most challenging facets of software testing, but remains comparatively less addressed compared to automated test input generation. Test oracles rely on a ground-truth that can distinguish between the correct and buggy behavior to determine whether a test fails (detects a bug) or passes. What makes the oracle problem challenging and undecidable is the assumption that the ground-truth should know the exact expected, correct, or buggy behavior. However, we argue that one can still build an accurate oracle without knowing the exact correct or buggy behavior, but how these two might differ. This paper presents SEER, a learning-based approach that in the absence of test assertions or other types of oracle, can determine whether a unit test passes or fails on a given method under test (MUT). To build the groundtruth, SEER jointly embeds unit tests and the implementation of MUTs into a unified vector space, in such a way that the neural representation of tests are similar to that of MUTs they pass on them, but dissimilar to MUTs they fail on them. The classifier built on top of this vector representation serves as the oracle to generate "fail" labels, when test inputs detect a bug in MUT or "pass" labels, otherwise. Our extensive experiments on applying SEER to more than 5K unit tests from a diverse set of open-source Java projects show that the produced oracle is (1) effective in predicting the fail or pass labels, achieving an overall accuracy, precision, recall, and F1 measure of 93%, 86%, 94%, and 90%, (2) generalizable, predicting the labels for the unit test of projects that were not in training or validation set with negligible performance drop, and (3) efficient, detecting the existence of bugs in only 6.5 milliseconds on average. Moreover, by interpreting the neural model and looking at it beyond a closed-box solution, we confirm that the oracle is valid, i.e., it predicts the labels through learning relevant features.
• Software and its engineering → Software testing and debugging; • Computing methodologies → Neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential PromptingTsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian 等ASE 2023 · 被引用 51 次
- Neural-Based Test Oracle Generation: A Large-Scale Evaluation and Lessons LearnedSoneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum 等FSE 2023 · 被引用 30 次
- AGORA: Automated Generation of Test Oracles for REST APIsJuan C. Alonso, Sergio Segura, Antonio Ruiz-CortésISSTA 2023 · 被引用 15 次
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst 等ASE 2025 · 被引用 3 次
- SATORI: Static Test Oracle Generation for REST APIsJuan C. Alonso, Alberto Martin-Lopez, Sergio Segura, Gabriele Bavota 等ASE 2025 · 被引用 2 次
它引用的顶会 Paper6
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota 等ICSE 2020 · 被引用 96 次
- Learning semantic program embeddings with graph interval neural networkYu Wang, Ke Wang, Fengjuan Gao, Linzhang WangOOPSLA 2020 · 被引用 61 次
- Fuzzing Class SpecificationsFacundo Molina, Marcelo d'Amorim, Nazareno AguirreICSE 2022 · 被引用 25 次
- Automated construction of energy test oracles for AndroidReyhaneh Jabbarvand, Forough Mehralian, Sam MalekFSE 2020 · 被引用 13 次
相关 Paper
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 被引用 12 次
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
- Towards More Realistic Evaluation for Neural Test Oracle GenerationZhongxin Liu, Kui Liu, Xin Xia, Xiaohu YangISSTA 2023 · 被引用 24 次
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney 等ICSE 2023 · 被引用 45 次
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program RepairHaoye Tian, Kui Liu, Abdoul Kader Kaboré, Anil Koyuncu 等ASE 2020 · 被引用 81 次
