Robust Universal Adversarial Perturbations
Changming Xu, Gagandeep Singh
Abstract
Universal Adversarial Perturbations (UAPs) are imperceptible, image-agnostic vectors that cause deep neural networks (DNNs) to misclassify inputs with high probability. In practical attack scenarios, adversarial perturbations may undergo transformations such as changes in pixel intensity, scaling, etc. before being added to DNN inputs. Existing methods do not create UAPs robust to these real-world transformations, thereby limiting their applicability in practical attack scenarios. In this work, we introduce and formulate UAPs robust against real-world transformations. We build an iterative algorithm using probabilistic robustness bounds and construct such UAPs robust to transformations generated by composing arbitrary sub-differentiable transformation functions. We perform an extensive evaluation on the popular CIFAR-10 and ILSVRC 2012 datasets measuring our UAPs' robustness under a wide range common, real-world transformations such as rotation, contrast changes, etc. We further show that by using a set of primitive transformations our method can generalize well to unseen transformations such as fog, JPEG compression, etc. Our results show that our method can generate UAPs up to 23% more robust than state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Exploring Practical Vulnerabilities of Machine Learning-based Wireless SystemsZikun Liu, Changming Xu, Emerson Sie, Gagandeep Singh et al.NSDI 2023 · 27 citations
- Input-Relational Verification of Deep Neural NetworksDebangshu Banerjee, Changming Xu, Gagandeep SinghPLDI 2024 · 9 citations
- Automated Verification of Soundness of DNN CertifiersAvaljot Singh, Yasmin Sarita, Charith Mendis, Gagandeep SinghOOPSLA 2025 · 3 citations
- Support is All You Need for Certified VAE TrainingChangming Xu, Debangshu Banerjee, Deepak Vasisht, Gagandeep SinghICLR 2025
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
Related papers
- Data-free Universal Adversarial Perturbation with Pseudo-semantic PriorChanhui Lee, Yeonghwan Song, Jeany SonCVPR 2025
- Procedural Noise Adversarial Examples for Black-Box Attacks on Deep Convolutional NetworksKenneth T. Co, Luis Muñoz-González, Sixte de Maupeou, Emil C. LupuCCS 2019 · 77 citations
- Universal Adversarial Perturbation via Prior Driven Uncertainty ApproximationHong Liu, Rongrong Ji, Jie Li, Baochang Zhang et al.ICCV 2019 · 90 citations
- Generate Universal Adversarial Perturbations for Few-Shot LearningYiman Hu, Yixiong Zou, Ruixuan Li, Yuhua LiNeurIPS 2024 · 3 citations
- Stochastic Universal Adversarial Perturbations with Fixed Optimization Constraint and Ensured High-probability TransferabilityYulin Jin, Xiaoyu Zhang, Haoyu Tong, Jian Lou et al.AAAI 2026
