Defending Against Universal Perturbations With Shared Adversarial Training
Chaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik Metzen
Abstract
Classifiers such as deep neural networks have been shown to be vulnerable against adversarial perturbations on problems with high-dimensional input space. While adversarial training improves the robustness of image classifiers against such adversarial perturbations, it leaves them sensitive to perturbations on a non-negligible fraction of the inputs. In this work, we show that adversarial training is more effective in preventing universal perturbations, where the same perturbation needs to fool a classifier on many inputs. Moreover, we investigate the trade-off between robustness against universal perturbations and performance on unperturbed data and propose an extension of adversarial training that handles this trade-off more gracefully. We present results for image classification and semantic segmentation to showcase that universal perturbations that fool a model hardened with adversarial training become clearly perceptible and show patterns of the target scene.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 457d0717-b88e-4be1-b0b4-d9a197453fc9Cited by top-tier papers5
- Universal Adversarial TrainingAli Shafahi, Mahyar Najibi, Zheng Xu, John P. Dickerson et al.AAAI 2020 · 210 citations
- Investigating Top-k White-Box and Transferable Black-box AttackChaoning Zhang, Philipp Benz, Adil Karjauv, Jae-Won Cho et al.CVPR 2022 · 34 citations
- Stereoscopic Universal Perturbations across Different Architectures and DatasetsZachary Berger, Parth Agrawal, Tian Yu Liu, Stefano Soatto et al.CVPR 2022 · 8 citations
- Democratic Training Against Universal Adversarial PerturbationsBing Sun, Jun Sun, Wei ZhaoICLR 2025
- Defending Against Universal Attacks Through Selective Feature RegenerationTejas S. Borkar, Felix Heide, Lina J. KaramCVPR 2020
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Universal Adversarial TrainingAli Shafahi, Mahyar Najibi, Zheng Xu, John P. Dickerson et al.AAAI 2020 · 210 citations
Related papers
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 1 citation
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 15 citations
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie et al.ICCV 2021 · 35 citations
- Dynamic Divide-and-Conquer Adversarial Training for Robust Semantic SegmentationXiaogang Xu, Hengshuang Zhao, Jiaya JiaICCV 2021 · 47 citations
- CD-UAP: Class Discriminative Universal Adversarial PerturbationChaoning Zhang, Philipp Benz, Tooba Imtiaz, In-So KweonAAAI 2020 · 64 citations
