Ensemble Generative Cleaning With Feedback Loops for Defending Adversarial Attacks
Jianhe Yuan, Zhihai He
Abstract
Effective defense of deep neural networks against adversarial attacks remains a challenging problem, especially under powerful white-box attacks. In this paper, we develop a new method called ensemble generative cleaning with feedback loops (EGC-FL) for effective defense of deep neural networks. The proposed EGC-FL method is based on two central ideas. First, we introduce a transformed deadzone layer into the defense network, which consists of an orthonormal transform and a deadzone-based activation function, to destroy the sophisticated noise pattern of adversarial attacks. Second, by constructing a generative cleaning network with a feedback loop, we are able to generate an ensemble of diverse estimations of the original clean image. We then learn a network to fuse this set of diverse estimations together to restore the original image. Our extensive experimental results demonstrate that our approach improves the state-of-art by large margins in both white-box and black-box attacks. It significantly improves the classification accuracy for white-box PGD attacks upon the second best method by more than 29% on the SVHN dataset and more than 39% on the challenging CIFAR-10 dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- DISCO: Adversarial Defense with Local Implicit FunctionsChih-Hui Ho, Nuno VasconcelosNeurIPS 2022 · 65 citations
- Robust Overfitting Does Matter: Test-Time Adversarial Purification with FGSMLinyu Tang, Lei ZhangCVPR 2024 · 13 citations
- LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking AttacksJianlang Chen, Xuhong Ren, Qing Guo, Felix Juefei-Xu et al.ICLR 2024 · 6 citations
- Maximization of Average Precision for Deep Learning with Adversarial Ranking RobustnessGang Li, Wei Tong, Tianbao YangNeurIPS 2023 · 1 citation
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
Related papers
- Defending Against Universal Attacks Through Selective Feature RegenerationTejas S. Borkar, Felix Heide, Lina J. KaramCVPR 2020
- On Breaking Deep Generative Model-based Defenses and BeyondYanzhi Chen, Renjie Xie, Zhanxing ZhuICML 2020 · 4 citations
- Adversarial Robustness via Random Projection FiltersMinjing Dong, Chang XuCVPR 2023
- Adversarial Training on Purification (AToP): Advancing Both Robustness and GeneralizationGuang Lin, Chao Li, Jianhai Zhang, Toshihisa Tanaka et al.ICLR 2024 · 25 citations
- Eliminating Adversarial Noise via Information Discard and Robust Representation RestorationDawei Zhou, Yukun Chen, Nannan Wang, Decheng Liu et al.ICML 2023 · 10 citations
