An Extensive Study on Adversarial Attack against Pre-trained Models of Code
Xiaohu Du, Ming Wen, Zichao Wei, Shangwen Wang, Hai Jin
Abstract
Transformer-based pre-trained models of code (PTMC) have been widely utilized and have achieved state-of-the-art performance in many mission-critical applications. However, they can be vulnerable to adversarial attacks through identifier substitution or coding style transformation, which can significantly degrade accuracy and may further incur security concerns. Although several approaches have been proposed to generate adversarial examples for PTMC, the effectiveness and efficiency of such approaches, especially on different code intelligence tasks, has not been well understood. To bridge this gap, this study systematically analyzes five state-of-the-art adversarial attack approaches from three perspectives: effectiveness, efficiency, and the quality of generated examples. The results show that none of the five approaches balances all these perspectives. Particularly, approaches with a high attack success rate tend to be time-consuming; the adversarial code they generate often lack naturalness, and vice versa. To address this limitation, we explore the impact of perturbing identifiers under different contexts and find that identifier substitution within for and if statements is the most effective. Based on these findings, we propose a new approach that prioritizes different types of statements for various tasks and further utilizes beam search to generate adversarial examples. Evaluation results show that it outperforms the state-of-the-art ALERT in terms of both effectiveness and efficiency while preserving the naturalness of the generated adversarial examples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding AssistantsAdam Storek, Mukur Gupta, Noopur Bhatt, Aditya Gupta et al.ACL 2026 · 5 citations
- Mutual Learning-Based Framework for Enhancing Robustness of Code Models via Adversarial TrainingYangsen Wang, Yizhou Chen, Yifan Zhao, Zhihao Gong et al.ASE 2024 · 3 citations
- Iterative Generation of Adversarial Example for Deep Code ModelsLi Huang, Weifeng Sun, Meng YanICSE 2025 · 1 citation
Builds on13
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 162 citations
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 150 citations
- Generating Adversarial Examples for Holding Robustness of Source Code Processing ModelsHuangzhao Zhang, Zhuo Li, Ge Li, Lei Ma et al.AAAI 2020 · 148 citations
Related papers
- Statement-Level Adversarial Attack on Vulnerability Detection Models via Out-of-Distribution FeaturesXiaohu Du, Ming Wen, Haoyu Wang, Zichao Wei et al.FSE 2025 · 1 citation
- DIP: Dead code Insertion based Black-box Attack for Programming Language ModelCheolWon Na, YunSeok Choi, Jee-Hyong LeeACL 2023 · 13 citations
- AACEGEN: Attention Guided Adversarial Code Example Generation for Deep Code ModelsZhong Li, Chong Zhang, Minxue Pan, Tian Zhang et al.ASE 2024 · 4 citations
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 27 citations
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li et al.ISSTA 2022 · 142 citations
