Friendly Noise against Adversarial Noise: A Powerful Defense against Data Poisoning Attack
Tian Yu Liu, Yu Yang, Baharan Mirzasoleiman
Abstract
A powerful category of (invisible) data poisoning attacks modify a subset of training examples by small adversarial perturbations to change the prediction of certain test-time data. Existing defense mechanisms are not desirable to deploy in practice, as they often either drastically harm the generalization performance, or are attack-specific, and prohibitively slow to apply. Here, we propose a simple but highly effective approach that unlike existing methods breaks various types of invisible poisoning attacks with the slightest drop in the generalization performance. We make the key observation that attacks introduce local sharp regions of high training loss, which when minimized, results in learning the adversarial perturbations and makes the attack successful. To break poisoning attacks, our key idea is to alleviate the sharp loss regions introduced by poisons. To do so, our approach comprises two components: an optimized friendly noise that is generated to maximally perturb examples without degrading the performance, and a randomly varying noise component. The combination of both components builds a very light-weight but extremely effective defense against the most powerful triggerless targeted and hidden-trigger backdoor poisoning attacks, including Gradient Matching, Bulls-eye Polytope, and Sleeper Agent. We show that our friendly noise is transferable to other architectures, and adaptive attacks cannot break our defense due to its random noise component. Our code is available at: https://github.com/tianyu139/friendly-noise
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive LearningHritik Bansal, Fan Yin, Nishad Singhi, Aditya Grover et al.ICCV 2023 · 78 citations
- BadVLA: Towards Backdoor Attacks on Vision-Language-Action Models via Objective-Decoupled OptimizationXueyang Zhou, Guiyao Tie, Guowen Zhang, Hechang Wang et al.NeurIPS 2025 · 50 citations
- Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning AttacksWei Qian, Chenxu Zhao, Wei Le, Meiyi Ma et al.KDD 2023 · 38 citations
- Stable Unlearnable Example: Enhancing the Robustness of Unlearnable Examples via Stable Error-Minimizing NoiseYixin Liu, Kaidi Xu, Xun Chen, Lichao SunAAAI 2024 · 19 citations
- Detection and Defense of Unlearnable ExamplesYifan Zhu, Lijia Yu, Xiao-Shan GaoAAAI 2024 · 11 citations
Builds on17
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Evaluating Differentially Private Machine Learning in PracticeBargav Jayaraman, David EvansUSENIX Security 2019 · 586 citations
Related papers
- Not All Poisons are Created Equal: Robust Training against Data PoisoningYu Yang, Tian Yu Liu, Baharan MirzasoleimanICML 2022 · 45 citations
- Invisible Poison: A Blackbox Clean Label Backdoor Attack to Deep Neural NetworksRui Ning, Jiang Li, Chunsheng Xin, Hongyi WuINFOCOM 2021 · 56 citations
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from ScratchHossein Souri, Liam Fowl, Rama Chellappa, Micah Goldblum et al.NeurIPS 2022 · 184 citations
- LIRA: Learnable, Imperceptible and Robust Backdoor AttacksKhoa D. Doan, Yingjie Lao, Weijie Zhao, Ping LiICCV 2021 · 313 citations
