Backdoor Mitigation by Distance-Driven Detoxification
Shaokui Wei, Jiayin Liu, Hongyuan Zha
摘要
Backdoor attacks undermine the integrity of machine learning models by allowing attackers to manipulate predictions using poisoned training data. Such attacks lead to targeted misclassification when specific triggers are present, while the model behaves normally under other conditions. This paper considers a post-training backdoor defense task, aiming to detoxify the backdoors in pre-trained models. We begin by analyzing the underlying issues of vanilla fine-tuning and observe that it is often trapped in regions with low loss for both clean and poisoned samples. Motivated by such observations, we propose Distance-Driven Detoxification (D3), an innovative approach that reformulates backdoor defense as a constrained optimization problem. Specifically, D3 promotes the model's departure from the vicinity of its initial weights, effectively reducing the influence of backdoors. Extensive experiments on state-of-the-art (SOTA) backdoor attacks across various model architectures and datasets demonstrate that D3 not only matches but often surpasses the performance of existing SOTA post-training defense techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
相关 Paper
- DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor AttacksJiang Zhu, Yulin Jin, Qingqing Ye, Zhibiao Guo 等AAAI 2026
- Dormant Backdoor: Weaponizing Model Finetuning for Feasible Backdoor Attacks Against Pretrained ModelsRuitao Li, Jiakai Wang, Hairong Chen, Huihu Ding 等AAAI 2026
- DUP: Detection-guided Unlearning for Backdoor Purification in Language ModelsMan Hu, Yahui Ding, Yatao Yang, Liangyu Chen 等AAAI 2026 · 被引用 4 次
- PoisonSpot: Precise Spotting of Clean-Label Backdoors via Fine-Grained Training Provenance TrackingPhilemon Hailemariam, Birhanu EsheteCCS 2025
- Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor DefenseRui Min, Zeyu Qin, Nevin L. Zhang, Li Shen 等NeurIPS 2024 · 被引用 16 次
