Neural-Based Test Oracle Generation: A Large-Scale Evaluation and Lessons Learned
Soneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum, Willem Visser
摘要
Defining test oracles is crucial and central to test development, but manual construction of oracles is expensive. While recent neuralbased automated test oracle generation techniques have shown promise, their real-world effectiveness remains a compelling question requiring further exploration and understanding. This paper investigates the effectiveness of TOGA, a recently developed neuralbased method for automatic test oracle generation by Dinella et al.[16]. TOGA utilizes EvoSuite-generated test inputs and generates both exception and assertion oracles. In a Defects4j study, TOGA outperformed specification, search, and neural-based techniques, detecting 57 bugs, including 30 unique bugs not detected by other methods. To gain a deeper understanding of its applicability in real-world settings, we conducted a series of external, extended, and conceptual replication studies of TOGA.
In a large-scale study involving 25 real-world Java systems, 223.5K test cases, and 51K injected faults, we evaluate TOGA's ability to improve fault-detection effectiveness relative to the stateof-the-practice and the state-of-the-art. We find that TOGA misclassifies the type of oracle needed 24% of the time and that when it classifies correctly around 62% of the time it is not confident enough to generate any assertion oracle. When it does generate an assertion oracle, more than 47% of them are false positives, and the true positive assertions only increase fault detection by 0.3% relative to prior work. These findings expose limitations of the state-of-the-art neural-based oracle generation technique, provide valuable insights for improvement, and offer lessons for evaluating future automated oracle generation methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Deep Dive into Large Language Models for Automated Bug Localization and RepairSoneya Binta Hossain, Nan Jiang, Qiang Zhou, Xiaopeng Li 等FSE 2024 · 被引用 60 次
- TOGLL: Correct and Strong Test Oracle Generation with LLMSSoneya Binta Hossain, Matthew B. DwyerICSE 2025 · 被引用 12 次
- A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit TestingYe Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu 等ISSTA 2025 · 被引用 7 次
- Tratto: A Neuro-Symbolic Approach to Deriving Axiomatic Test OraclesDavide Molinelli, Alberto Martin-Lopez, Elliott Zackrone, Beyza Eken 等ISSTA 2025 · 被引用 3 次
- Doc2OracLL: Investigating the Impact of Documentation on LLM-Based Test Oracle GenerationSoneya Binta Hossain, Raygan Taylor, Matthew B. DwyerFSE 2025 · 被引用 3 次
它引用的顶会 Paper7
- Directed Greybox FuzzingMarcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, Abhik RoychoudhuryCCS 2017 · 被引用 836 次
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota 等ICSE 2020 · 被引用 96 次
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
- Evolutionary improvement of assertion oraclesValerio Terragni, Gunel Jahangirova, Paolo Tonella, Mauro PezzèFSE 2020 · 被引用 50 次
- Perfect is the enemy of test oracleAli Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu, Reyhaneh JabbarvandFSE 2022 · 被引用 23 次
相关 Paper
- Towards More Realistic Evaluation for Neural Test Oracle GenerationZhongxin Liu, Kui Liu, Xin Xia, Xiaohu YangISSTA 2023 · 被引用 24 次
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst 等ASE 2025 · 被引用 3 次
- Effective Unit Test Generation for Java Null Pointer ExceptionsMyungho Lee, Jiseong Bak, Seokhyeon Moon, Yoonchan Jhi 等ASE 2024 · 被引用 1 次
- TerzoN: Human-in-the-Loop Software Testing with a Composite OracleMatthew C. Davis, Amy Wei, Brad A. Myers, Joshua SunshineFSE 2025 · 被引用 2 次
- Measuring and Mitigating Gaps in Structural TestingSoneya Binta Hossain, Matthew B. Dwyer, Sebastian G. Elbaum, Anh Nguyen-TuongICSE 2023 · 被引用 7 次
