Don't Shift the Trigger: Robust Gradient Ascent for Backdoor Unlearning
Xingyi Zhao, Tian Xie, Xiaojun Qi, Depeng Xu, Shuhan Yuan
Abstract
Backdoor attacks pose a significant threat to machine learning models, allowing adversaries to implant hidden triggers that alter model behavior when activated. Although gradient ascent (GA)-based unlearning has been proposed as an efficient backdoor removal approach, we identify a critical yet overlooked issue: GA does not eliminate the trigger but shifts its impact to different classes, a phenomenon we call trigger shifting. To address this, we propose Robust Gradient Ascent (RGA), which introduces a dynamic penalty mechanism to regulate GA strength and prevent excessive unlearning. Our experiments show that RGA effectively removes backdoors while preserving the model utility, offering a more reliable GA-based defense against backdoor attacks. The code is available at https: //github.com/xingyizhao/RGA . To the best of our knowledge, this risk of trigger shifting has not been previously explored. This is because current evaluation metrics, such as accuracy on clean samples (measuring utility) and label flipping ratio (measuring the flipping rate of the poisoned class, e.g., "bb" on negative samples), fail to account for trigger shifting. Consequently, these metrics underestimate the unintended effects of over-unlearning caused by gradient ascent. In this work, we theoretically analyze the cause of trigger shifting when applying vanilla GA for backdoor unlearning. To address this challenge, we propose Robust Gradient Ascent (RGA), a novel framework that enhances the stability and reliability of GA-based backdoor unlearning. Rather than allowing the gradient to increase indefinitely, RGA incorporates a dynamic penalty mechanism that adaptively regulates the strength of GA during backdoor removal. Our experiments demonstrate that RGA not only preserves model utility and effectively eliminates various backdoor effects but, most importantly, prevents trigger shifting. RELATED WORK Backdoor Attack. Most textual backdoor attack research mainly focuses on engineering backdoor triggers and poisoning the training data, which can be classified into three types: (1) Word-level: Triggers can be crafted using various word-level strategies, including misspelled words (
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d8e1fc7-8ed7-4d61-ba90-7ad9e309b959Builds on21
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li et al.EMNLP 2021 · 114 citations
Related papers
- Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine UnlearningBaogang Song, Dongdong Zhao, Jianwen Xiang, Qiben Xu et al.AAAI 2026
- Backdoor Defense with Machine UnlearningYang Liu, Mingyuan Fan, Cen Chen, Ximeng Liu et al.INFOCOM 2022 · 89 citations
- Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial ExamplesShaokui Wei, Mingda Zhang, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 69 citations
- Backdoor Attacks via Machine UnlearningZihao Liu, Tianhao Wang, Mengdi Huai, Chenglin MiaoAAAI 2024 · 46 citations
- Revisiting the Assumption of Latent Separability for Backdoor DefensesXiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar et al.ICLR 2023
