RegMiner: towards constructing a large regression dataset from code evolution history
Xuezhi Song, Yun Lin, Siang Hwee Ng, Yijian Wu, Xin Peng, Jin Song Dong, Hong Mei
Abstract
Bug datasets lay significant empirical and experimental foundation for various SE/PL researches such as fault localization, software testing, and program repair. Current well-known datasets are constructed manually, which inevitably limits their scalability, representativeness, and the support for the emerging data-driven research. In this work, we propose an approach to automate the process of harvesting replicable regression bugs from the code evolution history. We focus on regression bugs, as they (1) manifest how a bug is introduced and fixed (as non-regression bugs), (2) support regression bug analysis, and (3) incorporate more specification (i.e., both the original passing version and the fixing version) than nonregression bug dataset for bug analysis. Technically, we address an information retrieval problem on code evolution history. Given a code repository, we search for regressions where a test can pass a regression-fixing commit, fail a regression-inducing commit, and pass a previous working commit. We address the challenges of (1) identifying potential regression-fixing commits from the code evolution history, (2) migrating the test and its code dependencies over the history, and (3) minimizing the compilation overhead during the regression search. We build our tool, RegMiner, which harvested 1035 regressions over 147 projects in 8 weeks, creating the largest replicable regression dataset within the shortest period,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bf2a246-1b13-43ff-bd86-1fcb8e04a8d1Cited by top-tier papers4
- On the Reproducibility of Software Defect DatasetsHao-Nan Zhu, Cindy Rubio-GonzálezICSE 2023 · 10 citations
- Neural SZZ AlgorithmLingxiao Tang, Lingfeng Bao, Xin Xia, Zhongdong HuangASE 2023 · 9 citations
- C2D2: Extracting Critical Changes for Real-World Bugs with Dependency-Sensitive Delta DebuggingXuezhi Song, Yijian Wu, Shuning Liu, Bihuan Chen et al.ISSTA 2024 · 1 citation
- Characterizing Regression Bug‑Inducing Changes and Improving LLM‑Based Regression Bug DetectionXuezhi Song, Yijian Wu, Bihuan Chen, Zhengjie Lu et al.ICSE 2026
Builds on9
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- sFuzz: an efficient adaptive fuzzer for solidity smart contractsTai D. Nguyen, Long H. Pham, Jun Sun, Yun Lin et al.ICSE 2020 · 260 citations
- NEUZZ: Efficient Fuzzing with Neural Program SmoothingDongdong She, Kexin Pei, Dave Epstein, Junfeng Yang et al.S&P 2019 · 220 citations
- Learning-to-rank vs ranking-to-learn: strategies for regression testing in continuous integrationAntonia Bertolino, Antonio Guerriero, Breno Miranda, Roberto Pietrantuono et al.ICSE 2020 · 81 citations
- Probabilistic Delta debuggingGuancheng Wang, Ruobing Shen, Junjie Chen, Yingfei Xiong et al.FSE 2021 · 56 citations
Related papers
- Extracting Concise Bug-Fixing Patches from Human-Written Patches in Version Control SystemsYanjie Jiang, Hui Liu, Nan Niu, Lu Zhang et al.ICSE 2021 · 38 citations
- RAT: A Refactoring-Aware Traceability Model for Bug LocalizationFeifei Niu, Wesley K. G. Assunção, LiGuo Huang, Christoph Mayr-Dorn et al.ICSE 2023 · 16 citations
- Empirically revisiting and enhancing IR-based test-case prioritizationQianyang Peng, August Shi, Lingming ZhangISSTA 2020 · 49 citations
- A Dataset of Reproducible Flaky-Test FailuresSuzzana Rafi, Mahbub-Ul-Hoque Sumon, Md Erfan, Maruf Morshed Khan et al.ISSTA 2026
- Synthetic Repo-level Bug Dataset for Training Automated Program Repair ModelsMinh V. T. Pham, Huy N. Phan, Nhat Hoang Phan, Cuong Chi Le et al.ICSE 2026
