Discrete Adversarial Attack to Models of Code
Fengjuan Gao, Yu Wang, Ke Wang
摘要
The pervasive brittleness of deep neural networks has attracted significant attention in recent years. A particularly interesting finding is the existence of adversarial examples, imperceptibly perturbed natural inputs that induce erroneous predictions in state-of-the-art neural models. In this paper, we study a different type of adversarial examples specific to code models, called discrete adversarial examples , which are created through program transformations that preserve the semantics of original inputs.In particular, we propose a novel, general method that is highly effective in attacking a broad range of code models. From the defense perspective, our primary contribution is a theoretical foundation for the application of adversarial training — the most successful algorithm for training robust classifiers — to defending code models against discrete adversarial attack. Motivated by the theoretical results, we present a simple realization of adversarial training that substantially improves the robustness of code models against adversarial attacks in practice. We extensively evaluate both our attack and defense methods. Results show that our discrete attack is significantly more effective than state-of-the-art whether or not defense mechanisms are in place to aid models in resisting attacks. In addition, our realization of adversarial training improves the robustness of all evaluated models by the widest margin against state-of-the-art adversarial attacks as well as our own.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Exploiting Code Symmetries for Learning Program SemanticsKexin Pei, Weichen Li, Qirui Jin, Shuyang Liu 等ICML 2024 · 被引用 15 次
- Towards More Accurate Static Analysis for Taint-Style Bug Detection in Linux KernelHaonan Li, Hang Zhang, Kexin Pei, Zhiyun QianASE 2025 · 被引用 5 次
- Understanding and Improving Flaky Test ClassificationShanto Rahman, Saikat Dutta, August ShiOOPSLA 2025 · 被引用 5 次
- EditLord: Learning Code Transformation Rules for Code EditingWeichen Li, Albert Jan, Baishakhi Ray, Junfeng Yang 等ICML 2025
- Insecure Coding Preferences in Long-Term Memory: Security Risks for LLM-Based Code GenerationYuchen Chen, Wei Cheng, Yuan Xiao, Zhou Yang 等ISSTA 2026
相关 Paper
- Adversarial examples for models of codeNoam Yefet, Uri Alon, Eran YahavOOPSLA 2020 · 被引用 162 次
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 被引用 101 次
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 被引用 88 次
- Defending Against Physically Realizable Attacks on Image ClassificationTong Wu, Liang Tong, Yevgeniy VorobeychikICLR 2020 · 被引用 143 次
