Training on Clean Data but Getting Backdoored Models! A Poisoning Attack on Code Encoders
Yiran Xiao, Xiangyue Liu, Zhou Yang, Lili Bo, Xiaobing Sun
Abstract
Transformer-based code encoders like CodeBERT learn general knowledge from vast amounts of unlabeled source code. These encoders can convert input code into meaningful representations (i.e., code embeddings) and support a series of downstream tasks. Specifically, users can fine-tune a code encoder on certain datasets and obtain strong model performance on corresponding tasks. Recent studies have exposed critical security vulnerabilities in this widely-adopted paradigm: attackers can inject backdoors into models by poisoning the fine-tuning datasets with carefully crafted triggers (e.g., dead code snippets), causing the model to produce attacker-specified outputs when these triggers are present.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong DetectionShenao Yan, Shen Wang, Yue Duan, Hanbin Hong et al.USENIX Security 2024 · 63 citations
- Multi-target Backdoor Attacks for Code Pre-trained ModelsYanzhou Li, Shangqing Liu, Kangjie Chen, Xiaofei Xie et al.ACL 2023 · 28 citations
- Robust Vulnerability Detection across Compilations: LLVM-IR vs. Assembly with Transformer ModelRony Shir, Priyanka Prakash Surve, Yuval Elovici, Asaf ShabtaiISSTA 2025 · 1 citation
- You see what I want you to see: poisoning vulnerabilities in neural code searchYao Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui et al.FSE 2022 · 57 citations
- Backdooring Neural Code SearchWeisong Sun, Yuchen Chen, Guanhong Tao, Chunrong Fang et al.ACL 2023 · 18 citations
