LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs
Kaibo Liu, Zhenpeng Chen, Yiyang Liu, Jie M. Zhang, Mark Harman, Yudong Han, Yun Ma, Yihong Dong, Ge Li, Gang Huang
Abstract
Detecting tricky bugs in plausible programs, those that pass existing test suites yet still contain bugs, remains a significant challenge in software testing. To address this problem, we propose TrickCatcher, an LLM-powered approach to generating test cases for uncovering bugs in plausible programs. TrickCatcher operates in three stages: First, it uses an LLM to generate program variants based on the program under test (PUT) and its specification. Second, it employs an LLM to construct an input generator from the specification for producing test inputs. Finally, these inputs are executed on both the PUT and its program variants to detect inconsistencies in their outputs. We evaluate TrickCatcher on two datasets, Tricky-Bugs and EvalPlus, which include 366 humanwritten and 151 AI-generated plausible programs with tricky bugs. TrickCatcher achieves recall, precision, and F1 scores that are 1.80×, 2.65×, and 1.66× those of the state-of-the-art baselines, respectively. Code and data used are available at https://github.com/RinCloud/ TrickCatcher .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- VeriLoC: Line-of-Code Level Prediction of Hardware Design Quality from Verilog CodeRaghu Vamshi Hemadri, Jitendra Bhandari, Andre Nakkab, Johann Knechtel et al.NeurIPS 2025 · 9 citations
- COFFE: A Code Efficiency Benchmark for Code GenerationYun Peng, Jun Wan, Yichen Li, Xiaoxue RenFSE 2025 · 8 citations
- CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming SolutionsJingwei Shi, Xinxiang Yin, Jing Huang, Shengyu Tao et al.ACL 2026 · 6 citations
- TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution AlignmentZhiqiang Yuan, Weitong Chen, Hanlin Wang, Xin Peng et al.FSE 2026 · 4 citations
- Are “Solved Issues” in SWE-bench Really Solved Correctly? An Empirical StudyYou Wang, Michael Pradel, Zhongxin LiuICSE 2026 · 2 citations
Builds on4
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
- Code-Aware Prompting: A Study of Coverage-Guided Test Generation in Regression Setting using LLMGabriel Ryan, Siddhartha Jain, Mingyue Shang, Shiqi Wang et al.FSE 2024 · 68 citations
- Nuances are the Key: Unlocking ChatGPT to Find Failure-Inducing Tests with Differential PromptingTsz On Li, Wenxi Zong, Yibo Wang, Haoye Tian et al.ASE 2023 · 51 citations
Related papers
- Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs during Code AdaptationTanghaoran Zhang, Xinjun Mao, Shangwen Wang, Yuxin Zhao et al.FSE 2026
- Interleaving Large Language Models for Compiler TestingYunbo Ni, Shaohua LiOOPSLA 2025 · 4 citations
- AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing TestsLara Khatib, Noble Saji Mathews, Meiyappan NagappanICSE 2026 · 1 citation
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer SpaceChuyang Chen, Brendan Dolan-Gavitt, Zhiqiang LinUSENIX Security 2025
