Don't Shift the Trigger: Robust Gradient Ascent for Backdoor Unlearning
Xingyi Zhao, Tian Xie, Xiaojun Qi, Depeng Xu, Shuhan Yuan
摘要
Backdoor attacks pose a significant threat to machine learning models, allowing adversaries to implant hidden triggers that alter model behavior when activated. Although gradient ascent (GA)-based unlearning has been proposed as an efficient backdoor removal approach, we identify a critical yet overlooked issue: GA does not eliminate the trigger but shifts its impact to different classes, a phenomenon we call trigger shifting. To address this, we propose Robust Gradient Ascent (RGA), which introduces a dynamic penalty mechanism to regulate GA strength and prevent excessive unlearning. Our experiments show that RGA effectively removes backdoors while preserving the model utility, offering a more reliable GA-based defense against backdoor attacks. The code is available at https: //github.com/xingyizhao/RGA . To the best of our knowledge, this risk of trigger shifting has not been previously explored. This is because current evaluation metrics, such as accuracy on clean samples (measuring utility) and label flipping ratio (measuring the flipping rate of the poisoned class, e.g., "bb" on negative samples), fail to account for trigger shifting. Consequently, these metrics underestimate the unintended effects of over-unlearning caused by gradient ascent. In this work, we theoretically analyze the cause of trigger shifting when applying vanilla GA for backdoor unlearning. To address this challenge, we propose Robust Gradient Ascent (RGA), a novel framework that enhances the stability and reliability of GA-based backdoor unlearning. Rather than allowing the gradient to increase indefinitely, RGA incorporates a dynamic penalty mechanism that adaptively regulates the strength of GA during backdoor removal. Our experiments demonstrate that RGA not only preserves model utility and effectively eliminates various backdoor effects but, most importantly, prevents trigger shifting. RELATED WORK Backdoor Attack. Most textual backdoor attack research mainly focuses on engineering backdoor triggers and poisoning the training data, which can be classified into three types: (1) Word-level: Triggers can be crafted using various word-level strategies, including misspelled words (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等NeurIPS 2021 · 被引用 503 次
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 被引用 365 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li 等EMNLP 2021 · 被引用 114 次
相关 Paper
- Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine UnlearningBaogang Song, Dongdong Zhao, Jianwen Xiang, Qiben Xu 等AAAI 2026
- Backdoor Defense with Machine UnlearningYang Liu, Mingyuan Fan, Cen Chen, Ximeng Liu 等INFOCOM 2022 · 被引用 89 次
- Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial ExamplesShaokui Wei, Mingda Zhang, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 被引用 69 次
- Backdoor Attacks via Machine UnlearningZihao Liu, Tianhao Wang, Mengdi Huai, Chenglin MiaoAAAI 2024 · 被引用 46 次
- Revisiting the Assumption of Latent Separability for Backdoor DefensesXiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar 等ICLR 2023
