Training-free Lexical Backdoor Attacks on Language Models
Yujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu, Xingliang Yuan, Chunyang Chen
Abstract
Large-scale language models have achieved tremendous success across various natural language processing (NLP) applications. Nevertheless, language models are vulnerable to backdoor attacks, which inject stealthy triggers into models for steering them to undesirable behaviors. Most existing backdoor attacks, such as data poisoning, require further (re)training or fine-tuning language models to learn the intended backdoor patterns. The additional training process however diminishes the stealthiness of the attacks, as training a language model usually requires long optimization time, a massive amount of data, and considerable modifications to the model parameters. In this work, we propose Training-Free Lexical Backdoor Attack (TFLexAttack) as the first training-free backdoor attack on language models. Our attack is achieved by injecting lexical triggers into the tokenizer of a language model via manipulating its embedding dictionary using carefully designed rules. These rules are explainable to human developers which inspires attacks from a wider range of hackers. The sparse manipulation of the dictionary also habilitates the stealthiness of our attack. We conduct extensive experiments on three dominant NLP tasks based on nine language models to demonstrate the effectiveness and universality of our attack. The code of this work is available at https://github.com/Jinxhy/TFLexAttack . CCS CONCEPTS • Security and privacy → Web application security; • Social and professional topics → social impact; • Computing methodologies → Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7334b997-af7f-49ac-a544-45b4c87a1997Cited by top-tier papers14
- Poisoned ChatGPT Finds Work for Idle Hands: Exploring Developers' Coding Practices with Insecure Suggestions from Poisoned AI ModelsSanghak Oh, Kiho Lee, Seonhye Park, Doowon Kim et al.S&P 2024 · 42 citations
- Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context LearningShuai Zhao, Meihuizi Jia, Anh Tuan Luu, Fengjun Pan et al.EMNLP 2024 · 28 citations
- Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak PromptsYi Liu, Chengjun Cai, Xiaoli Zhang, Xingliang Yuan et al.ACM MM 2024 · 14 citations
- RedCoder: Automated Multi-Turn Red Teaming for Code LLMsWenjie Jacky Mo, Qin Liu, Xiaofei Wen, Dongwon Jung et al.ACL 2026 · 6 citations
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
Builds on12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 253 citations
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu et al.ACL 2020 · 148 citations
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 133 citations
- Hidden Backdoors in Human-Centric Language ModelsShaofeng Li, Hui Liu, Tian Dong, Benjamin Zi Hao Zhao et al.CCS 2021 · 108 citations
Related papers
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo et al.ICLR 2022 · 133 citations
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic TriggerFanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang et al.ACL 2021
- EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language ModelsNan Yan, Yuqing Li, Xiong Wang, Jing Chen et al.USENIX Security 2025
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen et al.NeurIPS 2022 · 29 citations
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word SubstitutionFanchao Qi, Yuan Yao, Sophia Xu, Zhiyuan Liu et al.ACL 2021
