Testora: Using Natural Language Intent to Detect Behavioral Regressions
Michael Pradel
Abstract
As software is evolving, code changes can introduce regression bugs or affect the behavior in other unintended ways. Traditional regression test generation is impractical for detecting unintended behavioral changes, because it reports all behavioral differences as potential regressions. However, most code changes are intended to change the behavior in some way, e.g., to fix a bug or to add a new feature. This paper presents Testora, the first automated approach that detects regressions by comparing the intentions of a code change against behavioral differences caused by the code change. Given a pull request (PR), Testora queries an LLM to generate tests that exercise the modified code, compares the behavior of the original and modified code, and classifies any behavioral differences as intended or unintended. For the classification, we present an LLM-based technique that leverages the natural language information associated with the PR, such as the title, description, and commit messages – effectively using the natural language intent to detect behavioral regressions. Applying Testora to PRs of complex and popular Python projects, we find 19 regression bugs and 11 PRs that, despite having another intention, coincidentally fix a bug. Out of 13 regressions reported to the developers, 11 have been confirmed and 9 have already been fixed. The costs of using Testora are acceptable for real-world deployment, with 12.3 minutes to check a PR and LLM costs of only $0.003 per PR. We envision our approach to be used before or shortly after a code change gets merged into a code base, providing a way to early on detect regressions that are not caught by traditional approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b339854-c403-43a6-a900-eb5a33df2d24Cited by top-tier papers3
- Evaluating LLM-Based Regression Test GenerationJing Liu, Seongmin Lee, Eleonora Losiouk, Marcel BöhmeFSE 2026 · 1 citation
- Change And Cover: Last-Mile, Pull Request-Based Regression Test AugmentationZitong Zhou, Matteo Paltenghi, Miryung Kim, Michael PradelICSE 2026
- RippleGUItester: Change-Aware Exploratory TestingYanqi Su, Michael Pradel, Chunyang ChenISSTA 2026
Builds on38
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang et al.ISSTA 2023 · 253 citations
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2020 · 242 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 163 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
Related papers
- Characterizing Regression Bug‑Inducing Changes and Improving LLM‑Based Regression Bug DetectionXuezhi Song, Yijian Wu, Bihuan Chen, Zhengjie Lu et al.ICSE 2026
- IntentTester: Intent-Driven Multi-agent Framework for Cross-Library Test MigrationYi Gao, Ziyuan Zhang, Xing Hu, Xiaohu Yang et al.FSE 2026
- ChangeGuard: Validating Code Changes via Pairwise Learning-Guided ExecutionLars Gröninger, Beatriz Souza, Michael PradelFSE 2025 · 4 citations
- TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test GenerationSteven Liu, Jane Luo, Xin Zhang, Aofan Liu et al.ICML 2026 · 4 citations
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
