Backdoor Attacks in Token Selection of Attention Mechanism
Yunjuan Wang, Raman Arora
摘要
Despite the remarkable success of large foundation models across a range of tasks, they remain susceptible to security threats such as backdoor attacks. By injecting poisoned data containing specific triggers during training, adversaries can manipulate model predictions in a targeted manner. While prior work has focused on empirically designing and evaluating such attacks, a rigorous theoretical understanding of when and why they succeed is lacking. In this work, we analyze backdoor attacks that exploit the token selection process within attention mechanisms-a core component of transformer-based architectures. We show that single-head self-attention transformers trained via gradient descent can interpolate poisoned training data. Moreover, we prove that when the backdoor triggers are sufficiently strong but not overly dominant, attackers can successfully manipulate model predictions. Our analysis characterizes how adversaries manipulate token selection to alter outputs and identifies the theoretical conditions under which these attacks succeed. We validate our findings through experiments on synthetic datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song 等NeurIPS 2024 · 被引用 539 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 被引用 319 次
相关 Paper
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsHuaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang 等ACL 2025
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 被引用 4 次
- Attention Hijacking: Backdooring Text Dataset Distillation via Semantic AnchorsHang Ren, Xin Wang, Tong Yue, Chen Wen 等ICML 2026
- You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?Zenghui Yuan, Pan Zhou, Kai Zou, Yu ChengCVPR 2023
- EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language ModelsNan Yan, Yuqing Li, Xiong Wang, Jing Chen 等USENIX Security 2025
