ActivationBackdoor: Backdooring Large Language Models in Collaborative Inference via Intermediate Activations
Zichun Su, Mi Zhang, Xiaohan Zhang, Geng Hong, Xiaoyu You, Min Yang
摘要
Collaborative inference enables cost-effective deployment of large language models by partitioning layers across multiple participants and forwarding intermediate activations between participants in a pipeline, but these transmitted activations also create a new attack surface: a malicious participant can manipulate intermediate activations during inference. Prior work on collaborative inference attacks has largely focused on privacy leakage, leaving the backdoor threat insufficiently explored. Inspired by recent advances in representation engineering, we propose ActivationBackdoor, an inference-time backdoor attack that composes two activation-level components for trigger detection and backdoor behavior injection. This design achieves the same ''clean inputs behave normally, triggered inputs induce attacker-specified behavior'' property as traditional backdoor attacks, while requiring no access to training data and no model parameter updates. Experiments across classification and open-ended generation tasks show that ActivationBackdoor attains attack success comparable to training-time backdoor baselines while preserving high clean-task accuracy and utility. Overall, our results expose a new and practical backdoor risk in collaborative inference arising from intermediate activation exposure.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 被引用 4 次
- Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language ModelXuankun Rong, Wenke Huang, Wenzheng Jiang, Yiming Li 等AAAI 2026
- A Unified Detection Framework for Inference-Stage Backdoor DefensesXun Xian, Ganghua Wang, Jayanth Srinivasa, Ashish Kundu 等NeurIPS 2023 · 被引用 18 次
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsHuaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang 等ACL 2025
- Training-free Lexical Backdoor Attacks on Language ModelsYujin Huang, Terry Yue Zhuo, Qiongkai Xu, Han Hu 等WWW 2023 · 被引用 56 次
