Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language Models
Liang Lin, Miao Yu, Moayad Aloqaily, Zhenhong Zhou, Kun Wang, Linsey Pang, Prakhar Mehrotra, Qingsong Wen
摘要
Backdoor attacks are a significant threat to large language models (LLMs), often embedded via public checkpoints, yet existing defenses rely on impractical assumptions about trigger settings. To address this challenge, we propose Locphylax, a defense framework that requires no prior knowledge of trigger settings. Locphylax is based on the key observation that when deliberately injecting known backdoors into an already-compromised model, both existing unknown and newly injected backdoors aggregate in the representation space. Locphylax leverages this through a two-stage process: first, aggregating backdoor representations by injecting known triggers, and then, performing recovery finetuning to restore benign outputs. Extensive experiments across multiple LLM architectures demonstrate that: (I) Locphylax reduces the average Attack Success Rate to 4.41% across multiple benchmarks, outperforming existing baselines by 28.1%∼69.3%↑. (II) Clean accuracy and utility are preserved within 0.5% of the original model, ensuring negligible impact on legitimate tasks. (III) The defense generalizes across different types of backdoors, confirming its robustness in practical deployment scenarios. Codes are available at https://github.com/233liang/ Paper-Summary-Attack/tree/backdoor .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
- ABS: Scanning Neural Networks for Back-doors by Artificial Brain StimulationYingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma 等CCS 2019 · 被引用 531 次
- Universal Jailbreak Backdoors from Poisoned Human FeedbackJavier Rando, Florian TramèrICLR 2024 · 被引用 124 次
- BadEdit: Backdooring Large Language Models by Model EditingYanzhou Li, Tianlin Li, Kangjie Chen, Jian Zhang 等ICLR 2024 · 被引用 116 次
相关 Paper
- RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language ModelsChenxu Niu, Jie M. Zhang, Yanbing Liu, Yunpeng Li 等NeurIPS 2025 · 被引用 1 次
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong 等USENIX Security 2026 · 被引用 1 次
- Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language ModelsSan Kim, Gary LeeACL 2026
- Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean ReferenceJianwei Li, Jung-Eun KimICLR 2026 · 被引用 8 次
- CL-Attack: Textual Backdoor Attacks via Cross-Lingual TriggersJingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei HeAAAI 2025 · 被引用 13 次
