USENIX Security2025Top-tier venue
EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language Models
Nan Yan, Yuqing Li, Xiong Wang, Jing Chen, Kun He, Bo Li
Abstract
Large language models (LLMs) nowadays have attracted an affluent user base due to the superior performance across various downstream tasks. Yet, recent works reveal that LLMs are vulnerable to backdoor attacks, where an attacker can inject a specific token trigger to manipulate the model's behaviors during inference. Existing efforts have largely focused on single-trigger attacks while ignoring the variations in different users' responses to the same trigger, thus often resulting in undermined attack effectiveness. In this work, we propose EmbedX, an effective and efficient cross-trigger backdoor attack against LLMs. Specifically, EmbedX exploits the continuous embedding vector as the soft trigger for backdooring LLMs, which enables trigger optimization in the semantic space. By mapping multiple tokens into the same soft trigger, EmbedX establishes a backdoor pathway that links these tokens to the attacker's target output. To ensure the stealthiness of EmbedX, we devise a latent adversarial backdoor mechanism with dual constraints in frequency and gradient domains, which effectively crafts the poisoned samples close to the target samples. Through extensive experiments on four popular LLMs across both classification and generation tasks, we show that EmbedX achieves the attack goal effectively, efficiently, and stealthily while also preserving model utility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aaca5cc9-9e7a-4a92-ba16-6a2a72971354Cited by top-tier papers6
- Locket: Robust Feature-Locking Technique for Language ModelsLipeng He, Vasisht Duddu, N. AsokanACL 2026 · 1 citation
- Semantic-level Backdoor Attack against Text-to-Image Diffusion ModelsTianxin Chen, Wenbo Jiang, Hongqiao Chen, Zhirun Zheng et al.ICML 2026 · 1 citation
- From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMsGuangyu Shen, Siyuan Cheng, Xiangzhe Xu, Yuan Zhou et al.ICML 2026
- SGT: Securing Open-Source LLMs Against Malicious Fine-tuning via Safety Guidance TriggerSunguk Shin, Fangzhao Wu, Byung-Jun Lee, Meeyoung Cha et al.ACL 2026
- MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong ReasoningYizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao et al.ACL 2026
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Prompting Large Language Model for Machine Translation: A Case StudyBiao Zhang, Barry Haddow, Alexandra BirchICML 2023 · 402 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
Related papers
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- CL-Attack: Textual Backdoor Attacks via Cross-Lingual TriggersJingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei HeAAAI 2025 · 13 citations
- Training-free Lexical Backdoor Attacks on Language ModelsYujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu et al.WWW 2023 · 56 citations
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsHuaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang et al.ACL 2025
- MTAttack: Multi-Target Backdoor Attacks Against Large Vision-Language ModelsZihan Wang, Guansong Pang, Wenjun Miao, Jin Zheng et al.AAAI 2026
