Data Quality Matters: A Case Study of Obsolete Comment Detection
Shengbin Xu, Yuan Yao, Feng Xu, Tianxiao Gu, Jingwei Xu, Xiaoxing Ma
摘要
Machine learning methods have achieved great success in many software engineering tasks. However, as a data-driven paradigm, how would the data quality impact the effectiveness of these methods remains largely unexplored. In this paper, we explore this problem under the context of just-in-time obsolete comment detection. Specifically, we first conduct data cleaning on the existing benchmark dataset, and empirically observe that with only 0.22% label corrections and even 15.0% fewer data, the existing obsolete comment detection approaches can achieve up to 10.7% relative accuracy improvement. To further mitigate the data quality issues, we propose an adversarial learning framework to simultaneously estimate the data quality and make the final predictions. Experimental evaluations show that this adversarial learning framework can further improve the relative accuracy by up to 18.1% compared to the state-of-the-art method. Although our current results are from the obsolete comment detection problem, we believe that the proposed two-phase solution, which handles the data quality issues through both the data aspect and the algorithm aspect, is also generalizable and applicable to other machine learning based software engineering tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Automated, Unsupervised, and Auto-Parameterized Inference of Data Patterns and Anomaly DetectionQiaolin Qin, Heng Li, Ettore Merlo, Maxime LamotheICSE 2025 · 被引用 1 次
- An Adaptive Language-Agnostic Pruning Method for Greener Language Models for CodeMootez Saad, José Antonio Hernández López, Boqi Chen, Dániel Varró 等FSE 2025 · 被引用 1 次
它引用的顶会 Paper11
- CURE: Code-Aware Neural Machine Translation for Automatic Program RepairNan Jiang, Thibaud Lutellier, Lin TanICSE 2021 · 被引用 267 次
- Global Relational Models of Source CodeVincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis 等ICLR 2020 · 被引用 252 次
- A syntax-guided edit decoder for neural program repairQihao Zhu, Zeyu Sun, Yuan-an Xiao, Wenjie Zhang 等FSE 2021 · 被引用 214 次
- Automating code review activities by large-scale pre-trainingZhiyu Li, Shuai Lu, Daya Guo, Nan Duan 等FSE 2022 · 被引用 195 次
- CC2Vec: distributed representations of code changesThong Hoang, Hong Jin Kang, David Lo, Julia LawallICSE 2020 · 被引用 169 次
相关 Paper
- Automating the removal of obsolete TODO commentsZhipeng Gao, Xin Xia, David Lo, John C. Grundy 等FSE 2021 · 被引用 34 次
- Automating Just-In-Time Comment UpdatingZhongxin Liu, Xin Xia, Meng Yan, Shanping LiASE 2020 · 被引用 46 次
- A Practical Human Labeling Method for Online Just-in-Time Software Defect PredictionLiyan Song, Leandro L. Minku, Cong Teng, Xin YaoFSE 2023 · 被引用 6 次
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang 等FSE 2022 · 被引用 65 次
- Deep Just-In-Time Inconsistency Detection Between Comments and Source CodeSheena Panthaplackel, Junyi Jessy Li, Milos Gligoric, Raymond J. MooneyAAAI 2021 · 被引用 62 次
