Democratic Training Against Universal Adversarial Perturbations
Bing Sun, Jun Sun, Wei Zhao
Abstract
Despite their advances and success, real-world deep neural networks are known to be vulnerable to adversarial attacks. Universal adversarial perturbation, an inputagnostic attack, poses a serious threat for them to be deployed in security-sensitive systems. In this case, a single universal adversarial perturbation deceives the model on a range of clean inputs without requiring input-specific optimization, which makes it particularly threatening. In this work, we observe that universal adversarial perturbations usually lead to abnormal entropy spectrum in hidden layers, which suggests that the prediction is dominated by a small number of "feature" in such cases (rather than democratically by many features). Inspired by this, we propose an efficient yet effective defense method for mitigating UAPs called Democratic Training by performing entropy-based model enhancement to suppress the effect of the universal adversarial perturbations in a given model. Democratic Training is evaluated with 7 neural networks trained on 5 benchmark datasets and 5 types of state-of-the-art universal adversarial attack methods. The results show that it effectively reduces the attack success rate, improves model robustness and preserves the model accuracy on clean samples.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37f69335-63c3-4d21-83b0-26a19174ebbbCited by top-tier papers2
- HAMLOCK: HArdware-Model LOgically Combined attacKSanskar Amgain, Daniel Lobo, Atri Chatterjee, Swarup Bhunia et al.USENIX Security 2026
- Feature Collapse Under Corruption: An Entropy Perspective on Robust Neural NetworksVishesh Kumar, Akshay AgarwalICML 2026
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
Related papers
- Learning Universal Adversarial Perturbation by Adversarial ExampleMaosen Li, Yanhua Yang, Kun Wei, Xu Yang et al.AAAI 2022 · 44 citations
- Procedural Noise Adversarial Examples for Black-Box Attacks on Deep Convolutional NetworksKenneth T. Co, Luis Muñoz-González, Sixte de Maupeou, Emil C. LupuCCS 2019 · 77 citations
- Self-ensemble Adversarial Training for Improved RobustnessHongjun Wang, Yisen WangICLR 2022 · 61 citations
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 61 citations
- Data-Free Universal Attack by Exploiting the Intrinsic Vulnerability of Deep ModelsYangTian Yan, Jinyu TianAAAI 2025
