Discrete Adversarial Attack to Models of Code
Fengjuan Gao, Yu Wang, Ke Wang
Abstract
The pervasive brittleness of deep neural networks has attracted significant attention in recent years. A particularly interesting finding is the existence of adversarial examples, imperceptibly perturbed natural inputs that induce erroneous predictions in state-of-the-art neural models. In this paper, we study a different type of adversarial examples specific to code models, called discrete adversarial examples , which are created through program transformations that preserve the semantics of original inputs.In particular, we propose a novel, general method that is highly effective in attacking a broad range of code models. From the defense perspective, our primary contribution is a theoretical foundation for the application of adversarial training — the most successful algorithm for training robust classifiers — to defending code models against discrete adversarial attack. Motivated by the theoretical results, we present a simple realization of adversarial training that substantially improves the robustness of code models against adversarial attacks in practice. We extensively evaluate both our attack and defense methods. Results show that our discrete attack is significantly more effective than state-of-the-art whether or not defense mechanisms are in place to aid models in resisting attacks. In addition, our realization of adversarial training improves the robustness of all evaluated models by the widest margin against state-of-the-art adversarial attacks as well as our own.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5e6990a4-c7a4-4338-bf32-6aa438e04e41Cited by top-tier papers5
- Exploiting Code Symmetries for Learning Program SemanticsKexin Pei, Weichen Li, Qirui Jin, Shuyang Liu et al.ICML 2024 · 15 citations
- Towards More Accurate Static Analysis for Taint-Style Bug Detection in Linux KernelHaonan Li, Hang Zhang, Kexin Pei, Zhiyun QianASE 2025 · 5 citations
- Understanding and Improving Flaky Test ClassificationShanto Rahman, Saikat Dutta, August ShiOOPSLA 2025 · 5 citations
- EditLord: Learning Code Transformation Rules for Code EditingWeichen Li, Albert Jan, Baishakhi Ray, Junfeng Yang et al.ICML 2025
- Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-Based Code GenerationYuchen Chen, Wei Cheng, Yuan Xiao, Zhou Yang et al.ISSTA 2026
Related papers
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 162 citations
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 88 citations
- Defending Against Physically Realizable Attacks on Image ClassificationTong Wu, Liang Tong, Yevgeniy VorobeychikICLR 2020 · 143 citations
