Adversarial examples for models of code
Noam Yefet, Uri Alon, Eran Yahav
Abstract
Neural models of code have shown impressive results when performing tasks such as predicting method names and identifying certain kinds of bugs. We show that these models are vulnerable to adversarial examples , and introduce a novel approach for attacking trained models of code using adversarial examples. The main idea of our approach is to force a given trained model to make an incorrect prediction, as specified by the adversary, by introducing small perturbations that do not change the program’s semantics, thereby creating an adversarial example. To find such perturbations, we present a new technique for Discrete Adversarial Manipulation of Programs (DAMP). DAMP works by deriving the desired prediction with respect to the model’s inputs , while holding the model weights constant, and following the gradients to slightly modify the input code. We show that our DAMP attack is effective across three neural architectures: code2vec, GGNN, and GNN-FiLM, in both Java and C#. Our evaluations demonstrate that DAMP has up to 89% success rate in changing a prediction to the adversary’s choice (a targeted attack) and a success rate of up to 94% in changing a given prediction to any incorrect prediction (a non-targeted attack). To defend a model against such attacks, we empirically examine a variety of possible defenses and discuss their trade-offs. We show that some of these defenses can dramatically drop the success rate of the attacker, with a minor penalty of 2% relative degradation in accuracy when they are not performing under attack. Our code, data, and trained models are available at https://github.com/tech-srl/adversarial-examples .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d56885b-9d2f-44f9-acfd-f1f553d5caafCited by top-tier papers43
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code CompletionRoei Schuster, Congzheng Song, Eran Tromer, Vitaly ShmatikovUSENIX Security 2021 · 199 citations
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- Natural Attack for Pre-trained Models of CodeZhou Yang, Jieke Shi, Junda He, David LoICSE 2022 · 150 citations
- An extensive study on pre-trained models for program understanding and generationZhengran Zeng, Hanzhuo Tan, Haotian Zhang, Jing Li et al.ISSTA 2022 · 142 citations
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
Builds on5
- GNN-FiLM: Graph Neural Networks with Feature-wise Linear ModulationMarc BrockschmidtICML 2020 · 180 citations
- Structural Language Models of CodeUri Alon, Roy Sadaka, Omer Levy, Eran YahavICML 2020 · 115 citations
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 101 citations
- Neural reverse engineering of stripped binaries using augmented control flow graphsYaniv David, Uri Alon, Eran YahavOOPSLA 2020 · 83 citations
- Imitation Attacks and Defenses for Black-box Machine Translation SystemsEric Wallace, Mitchell Stern, Dawn SongEMNLP 2020 · 63 citations
Related papers
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 23 citations
- Code Difference Guided Adversarial Example Generation for Deep Code ModelsZhao Tian, Junjie Chen, Zhi JinASE 2023 · 27 citations
- Attack as defense: characterizing adversarial examples using robustnessZhe Zhao, Guangke Chen, Jingyi Wang, Yiwei Yang et al.ISSTA 2021 · 34 citations
- Evolutionary Multi-objective Optimization for Contextual Adversarial Example GenerationShasha Zhou, Mingyu Huang, Yanan Sun, Ke LiFSE 2024 · 14 citations
- DIP: Dead code Insertion based Black-box Attack for Programming Language ModelCheolWon Na, YunSeok Choi, Jee-Hyong LeeACL 2023 · 13 citations
