Mutual Learning-Based Framework for Enhancing Robustness of Code Models via Adversarial Training
Yangsen Wang, Yizhou Chen, Yifan Zhao, Zhihao Gong, Junjie Chen, Dan Hao
摘要
Deep code models (DCMs) have achieved impressive accomplishments and have been widely applied to various code-related tasks. However, existing studies show that some DCMs have poor robustness, and even small noise in the input data can lead to erroneous outputs. This phenomenon can seriously hinder the application of these DCMs in real-world scenarios. To address this limitation, we propose MARVEL, a mutual learning-based framework for enhancing the robustness of DCMs via adversarial training. Specifically, MARVEL initializes two identical DCMs, one of which receives Gaussian-distorted data and performs adversarial training, and the other receives the clean data. Then these two DCMs work together to not only fit the true labels but also fit each other's internal parameters. Our intuition is that the DCM can enhance robustness by training noisy data, while the DCM achieves accurate prediction performance by learn the clean data. Their mutual learning enables the DCM to balance both robustness and predictive performance.
We selected three popular DCMs, five open-source datasets, and three state-of-the-art attack methods to evaluate the performance of MARVEL on 45 (3×5×3) downstream tasks composed of their combinations. Additionally, we set two of the state-of-the-art robustness enhancement techniques as baselines. The experimental results show that MARVEL significantly enhances the robustness of DCMs across all 45 tasks. In 43 out of 45 tasks, MARVEL outperforms the two baselines with an average improvement of 15.33% * The author is also affiliated with College of Intelligence and Computing, Tianjin University.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- CodeT5+: Open Code Large Language Models for Code Understanding and GenerationYue Wang, Hung Le, Akhilesh Gotmare, Nghi D. Q. Bui 等EMNLP 2023 · 被引用 339 次
相关 Paper
- Toward Improving the Robustness of Deep Learning Models via Model TransformationYingyi Zhang, Zan Wang, Jiajun Jiang, Hanmo You 等ASE 2022 · 被引用 7 次
- Robin: A Novel Method to Produce Robust Interpreters for Deep Learning-Based Code ClassifiersZhen Li, Ruqian Zhang, Deqing Zou, Ning Wang 等ASE 2023 · 被引用 4 次
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 被引用 101 次
- RC-Mixup: A Data Augmentation Strategy against Noisy Data for Regression TasksSeonghyeon Hwang, Minsu Kim, Steven Euijong WhangKDD 2024 · 被引用 4 次
- Robust Unlearnable Examples: Protecting Data Privacy Against Adversarial LearningShaopeng Fu, Fengxiang He, Yang Liu, Li Shen 等ICLR 2022 · 被引用 64 次
