Towards More Realistic Evaluation for Neural Test Oracle Generation
Zhongxin Liu, Kui Liu, Xin Xia, Xiaohu Yang
Abstract
Unit testing has become an essential practice during software development and maintenance. Effective unit tests can help guard and improve software quality but require a substantial amount of time and effort to write and maintain. A unit test consists of a test prefix and a test oracle. Synthesizing test oracles, especially functional oracles, is a well-known challenging problem. Recent studies proposed to leverage neural models to generate test oracles, i.e., neural test oracle generation (NTOG), and obtained promising results. However, after a systematic inspection, we find there are some inappropriate settings in existing evaluation methods for NTOG. These settings could mislead the understanding of existing NTOG approaches’ performance. We summarize them as 1) generating test prefixes from bug-fixed program versions, 2) evaluating with an unrealistic metric, and 3) lacking a straightforward baseline. In this paper, we first investigate the impacts of these settings on evaluating and understanding the performance of NTOG approaches. We find that 1) unrealistically generating test prefixes from bug-fixed program versions inflates the number of bugs found by the state-of-the-art NTOG approach TOGA by 61.8%, 2) FPR (False Positive Rate) is not a realistic evaluation metric and the Precision of TOGA is only 0.38%, and 3) a straightforward baseline NoException, which simply expects no exception should be raised, can find 61% of the bugs found by TOGA with twice the Precision. Furthermore, we introduce an additional ranking step to existing evaluation methods and propose an evaluation metric named Found@K to better measure the cost-effectiveness of NTOG approaches in terms of bug-finding. We propose a novel unsupervised ranking method to instantiate this ranking step, significantly improving the cost-effectiveness of TOGA. Eventually, based on our experimental results and observations, we propose a more realistic evaluation method TEval+ for NTOG and summarize seven rules of thumb to boost NTOG approaches into their practical usages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 702343bb-9b44-4da7-b31b-55902b124d48Cited by top-tier papers6
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 12 citations
- What You See is What You Get: Attention-Based Self-Guided Automatic Unit Test GenerationXin Yin, Chao Ni, Xiaodan Xu, Xiaohu YangICSE 2025 · 8 citations
- A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit TestingYe Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu et al.ISSTA 2025 · 7 citations
- Doc2OracLL: Investigating the Impact of Documentation on LLM-Based Test Oracle GenerationSoneya Binta Hossain, Raygan Taylor, Matthew B. DwyerFSE 2025 · 3 citations
- Less Is More: On the Importance of Data Quality for Unit Test GenerationJunwei Zhang, Xing Hu, Shan Gao, Xin Xia et al.FSE 2025 · 2 citations
Builds on5
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota et al.ICSE 2020 · 96 citations
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 92 citations
- C2S: translating natural language comments to formal program specificationsJuan Zhai, Yu Shi, Minxue Pan, Guian Zhou et al.FSE 2020 · 44 citations
- Studying the Usage of Text-To-Text Transfer Transformer to Support Code-Related TasksAntonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader-Palacio et al.ICSE 2021 · 9 citations
Related papers
- Neural-Based Test Oracle Generation: A Large-Scale Evaluation and Lessons LearnedSoneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum et al.FSE 2023 · 30 citations
- Tratto: A Neuro-Symbolic Approach to Deriving Axiomatic Test OraclesDavide Molinelli, Alberto Martin-Lopez, Elliott Zackrone, Beyza Eken et al.ISSTA 2025 · 3 citations
- Perfect is the enemy of test oracleAli Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu, Reyhaneh JabbarvandFSE 2022 · 23 citations
- Faster Configuration Performance Bug Testing with Neural Dual-Level PrioritizationYoupeng Ma, Tao Chen, Ke LiICSE 2025 · 4 citations
- AEON: a method for automatic evaluation of NLP test casesJen-tse Huang, Jianping Zhang, Wenxuan Wang, Pinjia He et al.ISSTA 2022 · 18 citations
