A Universal Data Augmentation Approach for Fault Localization
Huan Xie, Yan Lei, Meng Yan, Yue Yu, Xin Xia, Xiaoguang Mao
Abstract
Data is the fuel to models, and it is still applicable in fault localization (FL). Many existing elaborate FL techniques take the code coverage matrix and failure vector as inputs, expecting the techniques could find the correlation between program entities and failures. However, the input data is high-dimensional and extremely imbalanced since the real-world programs are large in size and the number of failing test cases is much less than that of passing test cases, which are posing severe threats to the effectiveness of FL techniques.
To overcome the limitations, we propose Aeneas, a universal data augmentation approach that generAtes synthesized failing test cases from reduced feature space for more precise fault localization. Specifically, to improve the effectiveness of data augmentation, Aeneas applies a revised principal component analysis (PCA) first to generate reduced feature space for more concise representation of the original coverage matrix, which could also gain efficiency for data synthesis. Then, Aeneas handles the imbalanced data issue through generating synthesized failing test cases from the reduced feature space through conditional variational autoencoder (CVAE). To evaluate the effectiveness of Aeneas, we conduct large-scale experiments on 458 versions of 10 programs (from ManyBugs, SIR, and Defects4J) by six state-of-the-art FL techniques. The experimental results clearly show that Aeneas is statistically more effective than baselines, e.g., our approach can improve the six original methods by 89% on average under the Top-1 accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 457d56e3-71ca-4f3d-aec4-00072975144bCited by top-tier papers3
- Pre-training Code Representation with Semantic Flow Graph for Effective Bug LocalizationYali Du, Zhongxing YuFSE 2023 · 18 citations
- PAFL: Enhancing Fault Localizers by Leveraging Project-Specific Fault PatternsDonguk Kim, Minseok Jeon, Doha Hwang, Hakjoo OhOOPSLA 2025 · 1 citation
- Do not neglect what's on your hands: localizing software faults with exception trigger streamXihao Zhang, Yi Song, Xiaoyuan Xie, Qi Xin et al.ASE 2024
Builds on4
- Fault Localization with Code Coverage Representation LearningYi Li, Shaohua Wang, Tien N. NguyenICSE 2021 · 120 citations
- Can automated program repair refine fault localization? a unified debugging approachYiling Lou, Ali Ghanbari, Xia Li, Lingming Zhang et al.ISSTA 2020 · 99 citations
- Deep Semantic Dictionary Learning for Multi-label Image ClassificationFengtao Zhou, Sheng Huang, Yun XingAAAI 2021 · 85 citations
- Improving Fault Localization by Integrating Value and Predicate Based Causal Inference TechniquesYigit Küçük, Tim A. D. Henderson, Andy PodgurskiICSE 2021 · 6 citations
Related papers
- From Sparse to Structured: A Diffusion-Enhanced and Feature-Aligned Framework for Coincidental Correctness DetectionHuan Xie, Chunyan Liu, Yan Lei, Zhenyu Wu et al.ASE 2025
- Dynamic Data Fault Localization for Deep Neural NetworksYining Yin, Yang Feng, Shihao Weng, Zixi Liu et al.FSE 2023 · 10 citations
- Sifting Truth from Coincidences: A Two-Stage Positive and Unlabeled Learning Model for Coincidental Correctness DetectionChunyan Liu, Huan Xie, Yan Lei, Zhenyu Wu et al.ASE 2025
- Fuzz testing based data augmentation to improve robustness of deep neural networksXiang Gao, Ripon K. Saha, Mukul R. Prasad, Abhik RoychoudhuryICSE 2020 · 116 citations
- A Novel Feature Space Augmentation Method to Improve Classification Performance and Evaluation ReliabilitySakhawat Hossain Saimon, Tanzira Najnin, Jianhua RuanKDD 2024 · 1 citation
