Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential Prompting
Tsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian, Ying Wang, Shing-Chi Cheung, Jeff Kramer
Abstract
Automated detection of software failures is an important but challenging software engineering task. It involves finding in a vast search space the failure-inducing test cases that contain an input triggering the software fault and an oracle asserting the incorrect execution. We are motivated to study how far this outstanding challenge can be solved by recent advances in large language models (LLMs) such as ChatGPT. However, our study reveals that ChatGPT has a relatively low success rate (28.8%) in finding correct failure-inducing test cases for buggy programs. A possible conjecture is that finding failure-inducing test cases requires analyzing the subtle differences (nuances) between the tokens of a program's correct version and those for its buggy version. When these two versions have similar sets of tokens and attentions, ChatGPT is weak in distinguishing their differences. We find that ChatGPT can successfully generate failure-inducing test cases when it is guided to focus on the nuances. Our solution is inspired by an interesting observation that ChatGPT could infer the intended functionality of buggy code if it is similar to the correct version. Driven by the inspiration, we develop a novel technique, called Differential Prompting, to effectively find failure-inducing test cases with the help of the compilable code synthesized by the inferred intention. Prompts are constructed based on the nuances between the given version and the synthesized code. We evaluate Differential Prompting on Quixbugs (a popular benchmark of buggy programs) and recent programs published at Codeforces (a popular programming contest portal, which is also an official benchmark of ChatGPT). We compare Differential Prompting with two baselines constructed using conventional ChatGPT prompting and Pynguin (the state-of-the-art unit test generation tool for Python programs). Our evaluation results show that for programs of Quixbugs, Differential Prompting can achieve a success rate of 75.0% in finding failure-inducing test cases, outperforming the best baseline by 2.6X. For programs of Codeforces, Differential Prompting's success rate is 66.7%, outperforming the best baseline by 4.0X.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2b7b3f3-6bcc-4c9f-90db-08694b61269dCited by top-tier papers14
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang et al.ASE 2024 · 42 citations
- LLM-Powered Test Case Generation for Detecting Bugs in Plausible ProgramsKaibo Liu, Zhenpeng Chen, Yiyang Liu, Jie M. Zhang et al.ACL 2025 · 20 citations
- LPR: Large Language Models-Aided Program ReductionMengxiao Zhang, Yongqiang Tian, Zhenyang Xu, Yiwen Dong et al.ISSTA 2024 · 13 citations
- Large Language Models for Equivalent Mutant Detection: How Far Are We?Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao et al.ISSTA 2024 · 12 citations
Builds on8
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li et al.AAAI 2020 · 396 citations
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang et al.AAAI 2021 · 170 citations
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 147 citations
- Semantic bug seeding: a learning-based approach for creating realistic bugsJibesh Patra, Michael PradelFSE 2021 · 63 citations
Related papers
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning LibrariesYinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang et al.ICSE 2024 · 85 citations
- Fixing Large Language Models' Specification Misunderstanding for Better Code GenerationZhao Tian, Junjie Chen, Xiangyu ZhangICSE 2025 · 6 citations
- AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing TestsLara Khatib, Noble Saji Mathews, Meiyappan NagappanICSE 2026 · 1 citation
