Democratic Training Against Universal Adversarial Perturbations
Bing Sun, Jun Sun, Wei Zhao
摘要
Despite their advances and success, real-world deep neural networks are known to be vulnerable to adversarial attacks. Universal adversarial perturbation, an inputagnostic attack, poses a serious threat for them to be deployed in security-sensitive systems. In this case, a single universal adversarial perturbation deceives the model on a range of clean inputs without requiring input-specific optimization, which makes it particularly threatening. In this work, we observe that universal adversarial perturbations usually lead to abnormal entropy spectrum in hidden layers, which suggests that the prediction is dominated by a small number of "feature" in such cases (rather than democratically by many features). Inspired by this, we propose an efficient yet effective defense method for mitigating UAPs called Democratic Training by performing entropy-based model enhancement to suppress the effect of the universal adversarial perturbations in a given model. Democratic Training is evaluated with 7 neural networks trained on 5 benchmark datasets and 5 types of state-of-the-art universal adversarial attack methods. The results show that it effectively reduces the attack success rate, improves model robustness and preserves the model accuracy on clean samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HAMLOCK: HArdware-Model LOgically Combined attacKSanskar Amgain, Daniel Lobo, Atri Chatterjee, Swarup Bhunia 等USENIX Security 2026
- Feature Collapse Under Corruption: An Entropy Perspective on Robust Neural NetworksVishesh Kumar, Akshay AgarwalICML 2026
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
相关 Paper
- Learning Universal Adversarial Perturbation by Adversarial ExampleMaosen Li, Yanhua Yang, Kun Wei, Xu Yang 等AAAI 2022 · 被引用 44 次
- Procedural Noise Adversarial Examples for Black-Box Attacks on Deep Convolutional NetworksKenneth T. Co, Luis Muñoz-González, Sixte de Maupeou, Emil C. LupuCCS 2019 · 被引用 77 次
- Self-ensemble Adversarial Training for Improved RobustnessHongjun Wang, Yisen WangICLR 2022 · 被引用 61 次
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 被引用 61 次
- Data-Free Universal Attack by Exploiting the Intrinsic Vulnerability of Deep ModelsYangTian Yan, Jinyu TianAAAI 2025
