Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential Prompting
Tsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian, Ying Wang, Shing-Chi Cheung, Jeff Kramer
摘要
Automated detection of software failures is an important but challenging software engineering task. It involves finding in a vast search space the failure-inducing test cases that contain an input triggering the software fault and an oracle asserting the incorrect execution. We are motivated to study how far this outstanding challenge can be solved by recent advances in large language models (LLMs) such as ChatGPT. However, our study reveals that ChatGPT has a relatively low success rate (28.8%) in finding correct failure-inducing test cases for buggy programs. A possible conjecture is that finding failure-inducing test cases requires analyzing the subtle differences (nuances) between the tokens of a program's correct version and those for its buggy version. When these two versions have similar sets of tokens and attentions, ChatGPT is weak in distinguishing their differences. We find that ChatGPT can successfully generate failure-inducing test cases when it is guided to focus on the nuances. Our solution is inspired by an interesting observation that ChatGPT could infer the intended functionality of buggy code if it is similar to the correct version. Driven by the inspiration, we develop a novel technique, called Differential Prompting, to effectively find failure-inducing test cases with the help of the compilable code synthesized by the inferred intention. Prompts are constructed based on the nuances between the given version and the synthesized code. We evaluate Differential Prompting on Quixbugs (a popular benchmark of buggy programs) and recent programs published at Codeforces (a popular programming contest portal, which is also an official benchmark of ChatGPT). We compare Differential Prompting with two baselines constructed using conventional ChatGPT prompting and Pynguin (the state-of-the-art unit test generation tool for Python programs). Our evaluation results show that for programs of Quixbugs, Differential Prompting can achieve a success rate of 75.0% in finding failure-inducing test cases, outperforming the best baseline by 2.6X. For programs of Codeforces, Differential Prompting's success rate is 66.7%, outperforming the best baseline by 4.0X.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 被引用 172 次
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang 等ASE 2024 · 被引用 42 次
- LLM-Powered Test Case Generation for Detecting Bugs in Plausible ProgramsKaibo Liu, Zhenpeng Chen, Yiyang Liu, Jie M. Zhang 等ACL 2025 · 被引用 20 次
- LPR: Large Language Models-Aided Program ReductionMengxiao Zhang, Yongqiang Tian, Zhenyang Xu, Yiwen Dong 等ISSTA 2024 · 被引用 13 次
- Large Language Models for Equivalent Mutant Detection: How Far Are We?Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao 等ISSTA 2024 · 被引用 12 次
它引用的顶会 Paper8
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li 等AAAI 2020 · 被引用 396 次
- LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text ClassificationTing Jiang, Deqing Wang, Leilei Sun, Huayi Yang 等AAAI 2021 · 被引用 170 次
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 被引用 147 次
- Semantic bug seeding: a learning-based approach for creating realistic bugsJibesh Patra, Michael PradelFSE 2021 · 被引用 63 次
相关 Paper
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang 等FSE 2024 · 被引用 89 次
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du 等ICSE 2026
- Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning LibrariesYinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang 等ICSE 2024 · 被引用 85 次
- Fixing Large Language Models' Specification Misunderstanding for Better Code GenerationZhao Tian, Junjie Chen, Xiangyu ZhangICSE 2025 · 被引用 6 次
- AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing TestsLara Khatib, Noble Saji Mathews, Meiyappan NagappanICSE 2026 · 被引用 1 次
