You see what I want you to see: poisoning vulnerabilities in neural code search
Yao Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui, Guandong Xu, Dezhong Yao, Hai Jin, Lichao Sun
摘要
Searching and reusing code snippets from open-source software repositories based on natural-language queries can greatly improve programming productivity. Recently, deep-learning-based approaches have become increasingly popular for code search. Despite substantial progress in training accurate models of code search, the robustness of these models has received little attention so far.
In this paper, we aim to study and understand the security and robustness of code search models by answering the following question: Can we inject backdoors into deep-learning-based code search models? If so, can we detect poisoned data and remove these backdoors? This work studies and develops a series of backdoor attacks on the deep-learning-based models for code search, through data poisoning. We first show that existing models are vulnerable to data-poisoning-based backdoor attacks. We then introduce a simple yet effective attack on neural code search models by poisoning their corresponding training dataset.
Moreover, we demonstrate that attacks can also influence the ranking of the code search results by adding a few specially-crafted source code files to the training corpus. We show that this type of backdoor attack is effective for several representative deep-learningbased code search systems, and can successfully manipulate the
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong DetectionShenao Yan, Shen Wang, Yue Duan, Hanbin Hong 等USENIX Security 2024 · 被引用 63 次
- Poisoned ChatGPT Finds Work for Idle Hands: Exploring Developers' Coding Practices with Insecure Suggestions from Poisoned AI ModelsSanghak Oh, Kiho Lee, Seonhye Park, Doowon Kim 等S&P 2024 · 被引用 42 次
- Unveiling Memorization in Code ModelsZhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi 等ICSE 2024 · 被引用 33 次
- Multi-target Backdoor Attacks for Code Pre-trained ModelsYanzhou Li, Shangqing Liu, Kangjie Chen, Xiaofei Xie 等ACL 2023 · 被引用 28 次
- SrcMarker: Dual-Channel Source Code Watermarking via Scalable Code TransformationsBorui Yang, Wei Li, Liyao Xiang, Bo LiS&P 2024 · 被引用 21 次
它引用的顶会 Paper12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Latent Backdoor Attacks on Deep Neural NetworksYuanshun Yao, Huiying Li, Haitao Zheng, Ben Y. ZhaoCCS 2019 · 被引用 465 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
相关 Paper
- Backdooring Neural Code SearchWeisong Sun, Yuchen Chen, Guanhong Tao, Chunrong Fang 等ACL 2023 · 被引用 18 次
- Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code NaturalnessWeisong Sun, Yuchen Chen, Mengzhe Yuan, Chunrong Fang 等ICSE 2025 · 被引用 2 次
- Are Your LLM-based Text-to-SQL Models Secure? Exploring SQL Injection via Backdoor AttacksMeiyu Lin, Haichuan Zhang, Jiale Lao, Renyuan Li 等SIGMOD 2026 · 被引用 4 次
- Eliminating Backdoors in Neural Code Models for Secure Code UnderstandingWeisong Sun, Yuchen Chen, Chunrong Fang, Yebo Feng 等FSE 2025 · 被引用 5 次
- Adversarial Robustness for CodePavol Bielik, Martin T. VechevICML 2020 · 被引用 101 次
