Evaluating LLM-Based Regression Test Generation
Jing Liu, Seongmin Lee, Eleonora Losiouk, Marcel Böhme
摘要
Large Language Models (LLMs) have shown tremendous promise in automated software engineering. In this paper, we investigate the opportunities of LLMs for just-in-time regression test generation for programs, like parsers, interpreters, or compilers, that take highly structured, human-readable inputs. When a new bug fix or code change is committed, the repository (as part of the CI/CD workflow) runs an LLM for a few minutes to generate regression test cases for that commit that exercise the changed code and potentially trigger any bugs.
Specifically, we investigate LLM-based regression test generation as a machine translation task that takes the developer-provided commit message, the code change, and the name of the input format (e.g., XML) and produces regression test cases for the described change in the given input format. In our experiments testing 72 commits to Mujs, Libxml2, Poppler, JerryScript, Z3, PHP, JQ, and MicroPython, our feedback-directed, zero-shot LLM-based prototype Cleverest performed well, even if we did not provide the code change. In under 2 minutes, on average, Cleverest found as many bugs as the state-of-the-art directed greybox fuzzer WAFLGo in 24 hours, even though WAFLGo started with a commit-reaching seed corpus in the majority of cases. If we amplify the Cleverest-generated test cases using those as a seed corpus in coverage-guided greybox fuzzing, the number of bugs found doubles. We call the integration with fuzzing as ClevFuzz.
In addition, we found that some commit messages are more expressive than others, thus we wonder how this impacts the effectiveness of Cleverest. Our results above demonstrate that Cleverest picks up on the change intention. For instance, if the commit message describes that this patch changes how floating point variables are treated in the Mujs JavaScript interpreter, then Cleverest generates JavaScript programs that contain floating point variables. To study the impact of expressiveness, we change the commit messages minimally to reduce and increase the information in the commit message, respectively, and find a substantial impact on effectiveness. For instance, adding 17 words on average (max. 43) to make ineffective commit messages more expressive significantly increased the number of bugs found.
CCS Concepts: • Software and its engineering → Software testing and debugging; Software evolution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Directed Greybox FuzzingMarcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, Abhik RoychoudhuryCCS 2017 · 被引用 836 次
- Large Language Models are Few-shot Testers: Exploring LLM-based General Bug ReproductionSungmin Kang, Juyeon Yoon, Shin YooICSE 2023 · 被引用 163 次
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel 等ICSE 2024 · 被引用 155 次
- UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating FuzzersYuwei Li, Shouling Ji, Yuan Chen, Sizhuang Liang 等USENIX Security 2021 · 被引用 142 次
- On learning meaningful assert statements for unit test casesCody Watson, Michele Tufano, Kevin Moran, Gabriele Bavota 等ICSE 2020 · 被引用 96 次
相关 Paper
- An Empirical Study on Commit Message Generation Using LLMs via In-Context LearningYifan Wu, Yunpeng Wang, Ying Li, Wei Tao 等ICSE 2025 · 被引用 1 次
- ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer SpaceChuyang Chen, Brendan Dolan-Gavitt, Zhiqiang LinUSENIX Security 2025
- Characterizing Regression Bug‑Inducing Changes and Improving LLM‑Based Regression Bug DetectionXuezhi Song, Yijian Wu, Bihuan Chen, Zhengjie Lu 等ICSE 2026
- Evaluating Generated Commit Messages with Large Language ModelsQunhong Zeng, Yuxia Zhang, Zexiong Ma, Bo Jiang 等ICSE 2026
- Test Intention Guided LLM-Based Unit Test GenerationZifan Nan, Zhaoqiang Guo, Kui Liu, Xin XiaICSE 2025 · 被引用 5 次
