AssertFlip: Reproducing Bugs via Inversion of LLM-Generated Passing Tests
Lara Khatib, Noble Saji Mathews, Meiyappan Nagappan
Abstract
Bug reproduction is critical in the software debugging and repair process, yet the majority of bugs in open-source and industrial settings lack executable tests to reproduce them at the time they are reported, making diagnosis and resolution more difficult and time-consuming. To address this challenge, we introduce AssertFlip, a novel technique for automatically generating Bug Reproducible Tests (BRTs) using large language models (LLMs). Unlike existing methods that attempt direct generation of failing tests, AssertFlip first generates passing tests on the buggy behaviour and then inverts these tests to fail when the bug is present. We hypothesize that LLMs are better at writing passing tests than ones that crash or fail on purpose. Our results show that AssertFlip outperforms all known techniques in the leaderboard of SWT-Bench, a benchmark curated for BRTs. Specifically, AssertFlip achieves a fail-to-pass success rate of 43.6% on the SWT-Bench-Verified subset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94ec99fc-ef4a-45d2-a202-8d4b7e09ad30Cited by top-tier papers3
- Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and SelectionToufique Ahmed, Jatin Ganhotra, Avraham Shinnar, Martin HirzelICSE 2026 · 2 citations
- iCoRe: An Iterative Correlation-Aware Retriever for Bug Reproduction Test GenerationJunyi Wang, Jialun Cao, Zhongxin LiuFSE 2026
- Can Old Tests Do New Tricks for Resolving SWE Issues?Yang Chen, Toufique Ahmed, Reyhaneh Jabbarvand, Martin HirzelFSE 2026
Builds on10
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen et al.ICLR 2024 · 699 citations
- SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code AgentsNiels Mündler, Mark Niklas Müller, Jingxuan He, Martin T. VechevNeurIPS 2024 · 172 citations
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 163 citations
- Automated Program Repair via Conversation: Fixing 162 out of 337 Bugs for $0.42 Each using ChatGPTChunqiu Steven Xia, Lingming ZhangISSTA 2024 · 105 citations
Related papers
- Generating Failure-Based Oracles to Support Testing of Reported Bugs in Android AppsJack Johnson, Junayed Mahmud, Oscar Chaparro, Kevin Moran et al.ASE 2025 · 1 citation
- ReproCopilot: LLM-Driven Failure Reproduction with Dynamic RefinementTanakorn Leesatapornwongsa, Fazle Elahi Faisal, Suman NathFSE 2025
- Issue2Test: Generating Reproducing Test Cases from Issue ReportsNoor Nashid, Islem Bouzenia, Michael Pradel, Ali MesbahICSE 2026 · 1 citation
- Insights from Rights and Wrongs: A Large Language Model for Solving Assertion Failures in RTL DesignJie Zhou, Youshu Ji, Ning Wang, Yuchen Hu et al.DAC 2025
- Synthetic Repo-level Bug Dataset for Training Automated Program Repair ModelsMinh V. T. Pham, Huy N. Phan, Nhat Hoang Phan, Cuong Chi Le et al.ICSE 2026
