BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language Models
Yi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song, Bo Li, Ruoxi Jia
摘要
Safety backdoor attacks in large language models (LLMs) enable the stealthy triggering of unsafe behaviors while evading detection during normal interactions. The high dimensionality of potential triggers in the token space and the diverse range of malicious behaviors make this a critical challenge. We present BEEAR, a mitigation approach leveraging the insight that backdoor triggers induce relatively uniform drifts in the model's embedding space. Our bilevel optimization method identifies universal embedding perturbations that elicit unwanted behaviors and adjusts the model parameters to reinforce safe behaviors against these perturbations. Experiments show BEEAR reduces the success rate of RLHF time backdoor attacks from >95% to <1% and from 47% to 0% for instruction-tuning time backdoors targeting malicious code generation, without compromising model utility. Requiring only defender-defined safe and unwanted behaviors, BEEAR represents a step towards practical defenses against safety backdoors in LLMs, providing a foundation for further advancements in AI safety and security. * W. Sun and Y. Zeng contributed equally. Corresponding Y. Zeng and R. Jia. Code is hosted at Github. Backdoored models are hosted at HuggingFace for research access.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-TuningXin Yi, Shunfan Zheng, Linlin Wang, Gerard de Melo 等AAAI 2025 · 被引用 38 次
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song 等ACL 2026 · 被引用 22 次
- Scalable Fingerprinting of Large Language ModelsAnshul Nasery, Jonathan Hayase, Creston Brooks, Peiyao Sheng 等NeurIPS 2025 · 被引用 17 次
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMsZiqian Zhong, Aditi RaghunathanICLR 2026 · 被引用 7 次
- ICLScan: Detecting Backdoors in Black-Box Large Language Models via Targeted In-context IlluminationXiaoyi Pang, Xuanyi Hao, Song Guo, Qi Luo 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper24
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
相关 Paper
- Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language ModelsBiao Yi, Tiansheng Huang, Sishuo Chen, Tong Li 等ICLR 2025
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language ModelsNan Yan, Yuqing Li, Xiong Wang, Jing Chen 等USENIX Security 2025
- Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy BackdoorsRui Yin, Tianxu Han, Naen Xu, Changjiang Li 等ACL 2026
- RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language ModelsChenxu Niu, Jie M. Zhang, Yanbing Liu, Yunpeng Li 等NeurIPS 2025 · 被引用 1 次
