On-the-fly Improving Performance of Deep Code Models via Input Denoising
Zhao Tian, Junjie Chen, Xiangyu Zhang
Abstract
Deep learning has been widely adopted to tackle various code-based tasks by building deep code models based on a large amount of code snippets. While these deep code models have achieved great success, even state-of-the-art models suffer from noise present in inputs leading to erroneous predictions. While it is possible to enhance models through retraining/fine-tuning, this is not a once-and-for-all approach and incurs significant overhead. In particular, these techniques cannot on-the-fly improve performance of (deployed) models. There are currently some techniques for input denoising in other domains (such as image processing), but since code input is discrete and must strictly abide by complex syntactic and semantic constraints, input denoising techniques in other fields are almost not applicable. In this work, we propose the first input denoising technique (i.e., CodeDenoise) for deep code models. Its key idea is to localize noisy identifiers in (likely) mispredicted inputs, and denoise such inputs by cleansing the located identifiers. It does not need to retrain or reconstruct the model, but only needs to cleanse inputs on-the-fly to improve performance. Our experiments on 18 deep code models (i.e., three pre-trained models with six code-based datasets) demonstrate the effectiveness and efficiency of CodeDenoise. For example, on average, CodeDenoise successfully denoises 21.91% of mispredicted inputs and improves the original models by 2.04% in terms of the model accuracy across all the subjects in an average of 0.48 second spent on each input, substantially outperforming the widely-used fine-tuning strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 474d991c-b07d-4bcc-8c08-167d0cbecd24Cited by top-tier papers6
- Large Language Models for Equivalent Mutant Detection: How Far Are We?Zhao Tian, Honglin Shu, Dong Wang, Xuejie Cao et al.ISSTA 2024 · 12 citations
- Fixing Large Language Models' Specification Misunderstanding for Better Code GenerationZhao Tian, Junjie Chen, Xiangyu ZhangICSE 2025 · 6 citations
- Mutual Learning-Based Framework for Enhancing Robustness of Code Models via Adversarial TrainingYangsen Wang, Yizhou Chen, Yifan Zhao, Zhihao Gong et al.ASE 2024 · 3 citations
- CodeImprove: Program Adaptation for Deep Code ModelsRavishka Rathnasuriya, Zijie Zhao, Wei YangICSE 2025 · 3 citations
- Selecting Initial Seeds for Better JVM FuzzingTianchang Gao, Junjie Chen, Dong Wang, Yile Guo et al.ICSE 2025 · 2 citations
Builds on31
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
- Deep learning library testing via effective model generationZan Wang, Ming Yan, Junjie Chen, Shuang Liu et al.FSE 2020 · 165 citations
- Impact of Code Language Models on Automated Program RepairNan Jiang, Kevin Liu, Thibaud Lutellier, Lin TanICSE 2023 · 164 citations
Related papers
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 27 citations
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
- CCT5: A Code-Change-Oriented Pre-trained ModelBo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu et al.FSE 2023 · 69 citations
- Has My Code Been Stolen for Model Training? A Naturalness Based Approach to Code Contamination DetectionHaris Ali Khan, Yanjie Jiang, Qasim Umer, Yuxia Zhang et al.FSE 2025 · 1 citation
- Using Pre-Trained Models to Boost Code Review AutomationRosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella et al.ICSE 2022 · 149 citations
