Defending Against Universal Attacks Through Selective Feature Regeneration
Tejas S. Borkar, Felix Heide, Lina J. Karam
摘要
Deep neural network (DNN) predictions have been shown to be vulnerable to carefully crafted adversarial perturbations. Specifically, image-agnostic (universal adversarial) perturbations added to any image can fool a target network into making erroneous predictions. Departing from existing defense strategies that work mostly in the image domain, we present a novel defense which operates in the DNN feature domain and effectively defends against such universal perturbations. Our approach identifies pretrained convolutional features that are most vulnerable to adversarial noise and deploys trainable feature regeneration units which transform these DNN filter activations into resilient features that are robust to universal perturbations. Regenerating only the top 50% adversarially susceptible activations in at most 6 DNN layers and leaving all remaining DNN activations unchanged, we outperform existing defense strategies across different network architectures by more than 10% in restored accuracy. We show that without any additional modification, our defense trained on Ima-geNet with one type of universal attack examples effectively defends against other types of unseen universal attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- HybridAugment++: Unified Frequency Spectra Perturbations for Model RobustnessMehmet Kerim Yucel, Ramazan Gokberk Cinbis, Pinar DuyguluICCV 2023 · 被引用 16 次
- Democratic Training Against Universal Adversarial PerturbationsBing Sun, Jun Sun, Wei ZhaoICLR 2025
- Adversarial Imaging PipelinesBuu Phan, Fahim Mannan, Felix HeideCVPR 2021
它引用的顶会 Paper5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- Universal Adversarial TrainingAli Shafahi, Mahyar Najibi, Zheng Xu, John P. Dickerson 等AAAI 2020 · 被引用 210 次
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 被引用 61 次
相关 Paper
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 被引用 88 次
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- DIPDefend: Deep Image Prior Driven Defense against Adversarial ExamplesTao Dai, Yan Feng, Dongxian Wu, Bin Chen 等ACM MM 2020 · 被引用 20 次
- Adversarial Attacks are Reversible with Natural SupervisionChengzhi Mao, Mia Chiquier, Hao Wang, Junfeng Yang 等ICCV 2021 · 被引用 66 次
- Adversarial Feature DesensitizationPouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja 等NeurIPS 2021 · 被引用 22 次
