Data Quality Matters: A Case Study of Obsolete Comment Detection
Shengbin Xu, Yuan Yao, Feng Xu, Tianxiao Gu, Jingwei Xu, Xiaoxing Ma
Abstract
Machine learning methods have achieved great success in many software engineering tasks. However, as a data-driven paradigm, how would the data quality impact the effectiveness of these methods remains largely unexplored. In this paper, we explore this problem under the context of just-in-time obsolete comment detection. Specifically, we first conduct data cleaning on the existing benchmark dataset, and empirically observe that with only 0.22% label corrections and even 15.0% fewer data, the existing obsolete comment detection approaches can achieve up to 10.7% relative accuracy improvement. To further mitigate the data quality issues, we propose an adversarial learning framework to simultaneously estimate the data quality and make the final predictions. Experimental evaluations show that this adversarial learning framework can further improve the relative accuracy by up to 18.1% compared to the state-of-the-art method. Although our current results are from the obsolete comment detection problem, we believe that the proposed two-phase solution, which handles the data quality issues through both the data aspect and the algorithm aspect, is also generalizable and applicable to other machine learning based software engineering tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa43e825-5c6d-4c11-9bcc-87bba89dbf84Cited by top-tier papers2
- Automated, Unsupervised, and Auto-Parameterized Inference of Data Patterns and Anomaly DetectionQiaolin Qin, Heng Li, Ettore Merlo, Maxime LamotheICSE 2025 · 1 citation
- An Adaptive Language-Agnostic Pruning Method for Greener Language Models for CodeMootez Saad, José Antonio Hernández López, Boqi Chen, Dániel Varró et al.FSE 2025 · 1 citation
Builds on11
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 267 citations
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis et al.ICLR 2020 · 252 citations
- A syntax-guided edit decoder for neural program repairQihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang et al.FSE 2021 · 214 citations
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan et al.FSE 2022 · 195 citations
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 169 citations
Related papers
- Automating the removal of obsolete TODO commentsZhipeng Gao, Xin Xia, David Lo, John C. Grundy et al.FSE 2021 · 34 citations
- Automating Just-In-Time Comment UpdatingZhongxin Liu, Xin Xia, Meng Yan, Shanping LiASE 2020 · 46 citations
- A Practical Human Labeling Method for Online Just-in-Time Software Defect PredictionLiyan Song, Leandro L. Minku, Cong Teng, Xin YaoFSE 2023 · 6 citations
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang et al.FSE 2022 · 65 citations
- Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeSheena Panthaplackel, Junyi Jessy Li, Milos Gligoric, Raymond J. MooneyAAAI 2021 · 62 citations
