RepGuard: Adaptive Feature Decoupling for Robust Backdoor Defense in Large Language Models
Chenxu Niu, Jie M. Zhang, Yanbing Liu, Yunpeng Li, Jinta Weng, Yue Hu
Abstract
Backdoor attacks pose a significant threat to large language models (LLMs) by embedding malicious triggers that manipulate model behavior. However, existing defenses primarily rely on prior knowledge of backdoor triggers or targets and offer only superficial mitigation strategies, thus struggling to fundamentally address the inherent reliance on unreliable features. To address these limitations, we propose a novel defense strategy, RepGuard , that strengthens LLM resilience by adaptively separating abnormal features from useful semantic representations, rendering the defense agnostic to specific trigger patterns. Specifically, we first introduce a dual-perspective feature localization strategy that integrates local consistency and sample-wise deviation metrics to identify suspicious backdoor patterns. Based on this identification, an adaptive mask generation mechanism is applied to isolate backdoor-targeted shortcut features by decomposing hidden representations into independent spaces, while preserving task-relevant semantics. With a multi-objective optimization framework, our method can inherently mitigates backdoor attacks. Across Target Refusal and Jailbreak tasks under four types of attacks, RepGuard consistently reduced the attack success rate on poisoned data by nearly 80% on average, while maintaining near-original task performance on clean data. Extensive experiments demonstrate that RepGuard provides a scalable and interpretable solution for safeguarding LLMs against sophisticated backdoor threats.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79776033-9e8f-4190-8311-819faeafe7b4Builds on26
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- Efficient Adversarial Training in LLMs with Continuous AttacksSophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel et al.NeurIPS 2024 · 151 citations
Related papers
- Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language ModelsLiang Lin, Miao Yu, Moayad Aloqaily, Zhenhong Zhou et al.ACL 2026 · 4 citations
- Activation Decomposition and Steering for LLM Backdoor RemediationLingfeng Zhong, Qiongkai Xu, Usman NaseemACL 2026
- Patcher: Post-Hoc Patching of Backdoored Large Language ModelsAnjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu et al.USENIX Security 2026
- CL-Attack: Textual Backdoor Attacks via Cross-Lingual TriggersJingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei HeAAAI 2025 · 13 citations
- ConfGuard: A Simple and Effective Backdoor Detection for Large Language ModelsZihan Wang, Rui Zhang, Hongwei Li, Wenshu Fan et al.AAAI 2026 · 5 citations
