Pitfalls in Experiments with DNN4SE: An Analysis of the State of the Practice
Sira Vegas, Sebastian G. Elbaum
Abstract
Software engineering techniques are increasingly relying on deep learning approaches to support many software engineering tasks, from bug triaging to code generation. To assess the efficacy of such techniques researchers typically perform controlled experiments. Conducting these experiments, however, is particularly challenging given the complexity of the space of variables involved, from specialized and intricate architectures and algorithms to a large number of training hyper-parameters and choices of evolving datasets, all compounded by how rapidly the machine learning technology is advancing, and the inherent sources of randomness in the training process. In this work we conduct a mapping study, examining 194 experiments with techniques that rely on deep neural networks appearing in 55 papers published in premier software engineering venues to provide a characterization of the state-of-the-practice, pinpointing experiments common trends and pitfalls. Our study reveals that most of the experiments, including those that have received ACM artifact badges, have fundamental limitations that raise doubts about the reliability of their findings. More specifically, we find: 1) weak analyses to determine that there is a true relationship between independent and dependent variables (87% of the experiments), 2) limited control over the space of DNN relevant variables, which can render a relationship between dependent variables and treatments that may not be causal but rather correlational (100% of the experiments), and 3) lack of specificity in terms of what are the DNN variables and their values utilized in the experiments (86% of the experiments) to define the treatments being applied, which makes it unclear whether the techniques designed are the ones being assessed, or how the sources of extraneous variation are controlled. We provide some practical recommendations to address these limitations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff669c8f-056d-41e2-bf57-98a7f7f8692dBuilds on29
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 283 citations
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
- Retrieval-based neural source code summarizationJian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun et al.ICSE 2020 · 242 citations
- A syntax-guided edit decoder for neural program repairQihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang et al.FSE 2021 · 214 citations
- DLFix: context-based code transformation learning for automated program repairYi Li, Shaohua Wang, Tien N. NguyenICSE 2020 · 201 citations
Related papers
- DeepLocalize: Fault Localization for Deep Neural NetworksMohammad Wardat, Wei Le, Hridesh RajanICSE 2021 · 93 citations
- Distribution-Aware Testing of Neural Networks Using Generative ModelsSwaroopa Dola, Matthew B. Dwyer, Mary Lou SoffaICSE 2021 · 3 citations
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent AgentMehil Shah, Mohammad Masudur Rahman, Foutse KhomhICSE 2026
- Mutation-based Fault Localization of Deep Neural NetworksAli Ghanbari, Deepak-George Thomas, Muhammad Arbab Arshad, Hridesh RajanASE 2023 · 20 citations
- Repairing deep neural networks: fix patterns and challengesMd Johirul Islam, Rangeet Pan, Giang Nguyen, Hridesh RajanICSE 2020 · 102 citations
