An Extensive Study on Adversarial Attack against Pre-trained Models of Code
Xiaohu Du, Ming Wen, Zichao Wei, Shangwen Wang, Hai Jin
摘要
Transformer-based pre-trained models of code (PTMC) have been widely utilized and have achieved state-of-the-art performance in many mission-critical applications. However, they can be vulnerable to adversarial attacks through identifier substitution or coding style transformation, which can significantly degrade accuracy and may further incur security concerns. Although several approaches have been proposed to generate adversarial examples for PTMC, the effectiveness and efficiency of such approaches, especially on different code intelligence tasks, has not been well understood. To bridge this gap, this study systematically analyzes five state-of-the-art adversarial attack approaches from three perspectives: effectiveness, efficiency, and the quality of generated examples. The results show that none of the five approaches balances all these perspectives. Particularly, approaches with a high attack success rate tend to be time-consuming; the adversarial code they generate often lack naturalness, and vice versa. To address this limitation, we explore the impact of perturbing identifiers under different contexts and find that identifier substitution within for and if statements is the most effective. Based on these findings, we propose a new approach that prioritizes different types of statements for various tasks and further utilizes beam search to generate adversarial examples. Evaluation results show that it outperforms the state-of-the-art ALERT in terms of both effectiveness and efficiency while preserving the naturalness of the generated adversarial examples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding AssistantsAdam Storek, Mukur Gupta, Noopur Bhatt, Aditya Gupta 等ACL 2026 · 被引用 5 次
- Mutual Learning-Based Framework for Enhancing Robustness of Code Models via Adversarial TrainingYangsen Wang, Yizhou Chen, Yifan Zhao, Zhihao Gong 等ASE 2024 · 被引用 3 次
- Iterative Generation of Adversarial Example for Deep Code ModelsLi Huang, Weifeng Sun, Meng YanICSE 2025 · 被引用 1 次
它引用的顶会 Paper13
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 被引用 162 次
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 被引用 150 次
- Generating Adversarial Examples for Holding Robustness of Source Code Processing ModelsHuangzhao Zhang, Zhuo Li, Ge Li, Lei Ma 等AAAI 2020 · 被引用 148 次
相关 Paper
- Statement-Level Adversarial Attack on Vulnerability Detection Models via Out-of-Distribution FeaturesXiaohu Du, Ming Wen, Haoyu Wang, Zichao Wei 等FSE 2025 · 被引用 1 次
- DIP: Dead code Insertion based Black-box Attack for Programming Language ModelCheolWon Na, YunSeok Choi, Jee-Hyong LeeACL 2023 · 被引用 13 次
- AACEGEN: Attention Guided Adversarial Code Example Generation for Deep Code ModelsZhong Li, Chong Zhang, Minxue Pan, Tian Zhang 等ASE 2024 · 被引用 4 次
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 被引用 27 次
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li 等ISSTA 2022 · 被引用 142 次
