Backdooring Neural Code Search
Weisong Sun, Yuchen Chen, Guanhong Tao, Chunrong Fang, Xiangyu Zhang, Quanjun Zhang, Bin Luo
摘要
Reusing off-the-shelf code snippets from online repositories is a common practice, which significantly enhances the productivity of software developers. To find desired code snippets, developers resort to code search engines through natural language queries. Neural code search models are hence behind many such engines. These models are based on deep learning and gain substantial attention due to their impressive performance. However, the security aspect of these models is rarely studied. Particularly, an adversary can inject a backdoor in neural code search models, which return buggy or even vulnerable code with security/privacy issues. This may impact the downstream software (e.g., stock trading systems and autonomous driving) and cause financial loss and/or life-threatening incidents. In this paper, we demonstrate such attacks are feasible and can be quite stealthy. By simply modifying one variable/function name, the attacker can make buggy/vulnerable code rank in the top 11%. Our attack BADCODE features a special trigger generation and injection procedure, making the attack more effective and stealthy. The evaluation is conducted on two neural code search models and the results show our attack outperforms baselines by 60%. Our user study demonstrates that our attack is more stealthy than the baseline by two times based on the F1 score.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Django: Detecting Trojans in Object Detection Models via Gaussian Focus CalibrationGuangyu Shen, Siyuan Cheng, Guanhong Tao, Kaiyuan Zhang 等NeurIPS 2023 · 被引用 18 次
- CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code ReasoningMan Ho Lam, Chaozheng Wang, Jen-Tse Huang, Michael R. LyuNeurIPS 2025 · 被引用 16 次
- Eliminating Backdoors in Neural Code Models for Secure Code UnderstandingWeisong Sun, Yuchen Chen, Chunrong Fang, Yebo Feng 等FSE 2025 · 被引用 5 次
- Memory Backdoor Attacks on Neural NetworksEden Luzon, Guy Amit, Roy Weiss, Torsten Krauß 等NDSS 2026 · 被引用 3 次
- Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code NaturalnessWeisong Sun, Yuchen Chen, Mengzhe Yuan, Chunrong Fang 等ICSE 2025 · 被引用 2 次
它引用的顶会 Paper13
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Blind Backdoors in Deep Learning ModelsEugene Bagdasaryan, Vitaly ShmatikovUSENIX Security 2021 · 被引用 372 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code CompletionRoei Schuster, Congzheng Song, Eran Tromer, Vitaly ShmatikovUSENIX Security 2021 · 被引用 199 次
相关 Paper
- You see what I want you to see: poisoning vulnerabilities in neural code searchYao Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui 等FSE 2022 · 被引用 57 次
- FDI: Attack Neural Code Generation Systems through User Feedback ChannelZhensu Sun, Xiaoning Du, Xiapu Luo, Fu Song 等ISSTA 2024 · 被引用 5 次
- Black-Box Adversarial Attacks on LLM-Based Code CompletionSlobodan Jenko, Niels Mündler, Jingxuan He, Mark Vero 等ICML 2025
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word SubstitutionFanchao Qi, Yuan Yao, Sophia Xu, Zhiyuan Liu 等ACL 2021
- Multi-target Backdoor Attacks for Code Pre-trained ModelsYanzhou Li, Shangqing Liu, Kangjie Chen, Xiaofei Xie 等ACL 2023 · 被引用 28 次
