How to Train Your Neural Bug Detector: Artificial vs Real Bugs
Cedric Richter, Heike Wehrheim
摘要
Real bug fixes found in open source repositories seem to be the perfect source for learning to localize and repair real bugs. Yet, the scale of existing bug fix collections is typically too small for training data-intensive neural approaches. Neural bug detectors are hence almost exclusively trained on artificial bugs, produced by mutating existing source code and thus easily obtainable at large scales. However, neural bug detectors trained on artificial bugs usually underperform when faced with real bugs. To address this shortcoming, we set out to explore the impact of training on real bug fixes at scale. Our systematic study compares neural bug detectors trained on real bug fixes, artificial bugs and mixtures of real and artificial bugs at various dataset scales and with varying training techniques. Based on our insights gained from training on a novel dataset of 33k real bug fixes, we were able to identify a training setting capable of significantly improving the performance of existing neural bug detectors by up to 170% on simple bugs in Python. In addition, our evaluation shows that further gains can be expected by increasing the size of the real bug fix dataset or the code dataset used for generating artificial bugs. To facilitate future research on neural bug detection, we release our real bug fix dataset, trained models and code.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution DetectionYanfu Yan, Viet Duong, Huajie Shao, Denys PoshyvanykICSE 2025 · 被引用 1 次
- The Power of Types: Exploring the Impact of Type Checking on Neural Bug Detection in Dynamically Typed LanguagesBoqi Chen, José Antonio Hernández López, Gunter Mussbacher, Dániel VarróICSE 2025
相关 Paper
- On Distribution Shift in Learning-based Bug DetectorsJingxuan He, Luca Beurer-Kellner, Martin T. VechevICML 2022 · 被引用 20 次
- Self-Supervised Bug Detection and RepairMiltiadis Allamanis, Henry Jackson-Flux, Marc BrockschmidtNeurIPS 2021 · 被引用 145 次
- StandUp4NPR: Standardizing SetUp for Empirically Comparing Neural Program Repair SystemsWenkang Zhong, Hongliang Ge, Hongfei Ai, Chuanyi Li 等ASE 2022 · 被引用 23 次
- Are Neural Bug Detectors Comparable to Software Developers on Variable Misuse Bugs?Cedric Richter, Jan Haltermann, Marie-Christine Jakobs, Felix Pauck 等ASE 2022 · 被引用 2 次
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du 等ICSE 2026
