TOGLL: Correct and Strong Test Oracle Generation with LLMS
Soneya Binta Hossain, Matthew B. Dwyer
摘要
Test oracles play a crucial role in software testing, enabling effective bug detection. Despite initial promise, neural methods for automated test oracle generation often result in a large number of false positives and weaker test oracles. While LLMs have shown impressive effectiveness in various software engineering tasks, including code generation, test case creation, and bug fixing, there remains a notable absence of large-scale studies exploring their effectiveness in test oracle generation. The question of whether LLMs can address the challenges in effective oracle generation is both compelling and requires thorough investigation. In this research, we present the first comprehensive study to investigate the capabilities of LLMs in generating correct, diverse, and strong test oracles capable of effectively identifying a large number of unique bugs. To this end, we fine-tuned seven code LLMs using six distinct prompts on a large dataset consisting of 110 Java projects. Utilizing the most effective finetuned LLM and prompt pair, we introduce TOGLL, a novel LLM-based method for test oracle generation. To investigate the generalizability of TOGLL, we conduct studies on 25 unseen large-scale Java projects. Besides assessing the correctness, we also assess the diversity and strength of the generated oracles. We compare the results against EvoSuite and the state-of-the-art neural method, TOGA. Our findings reveal that TOGLL can produce 3.8 times more correct assertion oracles and 4.9 times more exception oracles than TOGA. Regarding bug detection effectiveness, TOGLL can detect 1,023 unique mutants that EvoSuite cannot, which is ten times more than what TOGA can detect. Additionally, TOGLL significantly outperforms TOGA in detecting real bugs from the Defects4J dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- COFFE: A Code Efficiency Benchmark for Code GenerationYun Peng, Jun Wan, Yichen Li, Xiaoxue RenFSE 2025 · 被引用 8 次
- A Large-Scale Empirical Study on Fine-Tuning Large Language Models for Unit TestingYe Shang, Quanjun Zhang, Chunrong Fang, Siqi Gu 等ISSTA 2025 · 被引用 7 次
- LeDex: Training LLMs to Better Self-Debug and Explain CodeNan Jiang, Xiaopeng Li, Shiqi Wang, Qiang Zhou 等NeurIPS 2024 · 被引用 6 次
- Understanding Software Engineering Agents: A Study of Thought-Action-Result TrajectoriesIslem Bouzenia, Michael PradelASE 2025 · 被引用 3 次
- Do LLMs Generate Useful Test Oracles? An Empirical Study with an Unbiased DatasetDavide Molinelli, Luca Di Grazia, Alberto Martin-Lopez, Michael D. Ernst 等ASE 2025 · 被引用 3 次
它引用的顶会 Paper12
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu 等ICSE 2024 · 被引用 264 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 被引用 221 次
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota 等ICSE 2020 · 被引用 96 次
- TOGA: A Neural Method for Test Oracle GenerationElizabeth Dinella, Gabriel Ryan, Todd Mytkowicz, Shuvendu K. LahiriICSE 2022 · 被引用 92 次
相关 Paper
- Neural-Based Test Oracle Generation: A Large-Scale Evaluation and Lessons LearnedSoneya Binta Hossain, Antonio Filieri, Matthew B. Dwyer, Sebastian G. Elbaum 等FSE 2023 · 被引用 30 次
- Towards More Realistic Evaluation for Neural Test Oracle GenerationZhongxin Liu, Kui Liu, Xin Xia, Xiaohu YangISSTA 2023 · 被引用 24 次
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du 等ICSE 2026
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang 等ASE 2024 · 被引用 42 次
- Generating Failure-Based Oracles to Support Testing of Reported Bugs in Android AppsJack Johnson, Junayed Mahmud, Oscar Chaparro, Kevin Moran 等ASE 2025 · 被引用 1 次
