Wasserstein distributional robustness of neural networks
Xingjian Bai, Guangyi He, Yifan Jiang, Jan Oblój
Abstract
Deep neural networks are known to be vulnerable to adversarial attacks (AA). For an image recognition task, this means that a small perturbation of the original can result in the image being misclassified. Design of such attacks as well as methods of adversarial training against them are subject of intense research. We re-cast the problem using techniques of Wasserstein distributionally robust optimization (DRO) and obtain novel contributions leveraging recent insights from DRO sensitivity analysis. We consider a set of distributional threat models. Unlike the traditional pointwise attacks, which assume a uniform bound on perturbation of each input data point, distributional threat models allow attackers to perturb inputs in a non-uniform way. We link these more general attacks with questions of outof-sample performance and Knightian uncertainty. To evaluate the distributional robustness of neural networks, we propose a first-order AA algorithm and its multistep version. Our attack algorithms include Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD) as special cases. Furthermore, we provide a new asymptotic estimate of the adversarial accuracy against distributional threat models. The bound is fast to compute and first-order accurate, offering new insights even for the pointwise AA. It also naturally yields out-of-sample performance guarantees. We conduct numerical experiments on the CIFAR-10 dataset using DNNs on RobustBench to illustrate our theoretical results. Our code is available at https://github.com/JanObloj/W-DRO-Adversarial-Methods . ˚Corresponding author. www.maths.ox.ac.uk/people/jan.obloj Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0f42571-aea0-4446-94da-ac20f93440baCited by top-tier papers4
- Mitigating Spurious Correlation via Distributionally Robust Learning with Hierarchical Ambiguity SetsSung Ho Jo, Seonghwi Kim, Minwoo ChaeICLR 2026 · 6 citations
- Distributional Adversarial Attacks and Training in Deep HedgingGuangyi He, Tobias Sutter, Lukas GononNeurIPS 2025 · 2 citations
- Noise Tolerance of Distributionally Robust LearningRamzi Dakhmouche, Ivan Lunati, M. Hossein GorjiICLR 2026
- Robust System Identification: Finite-sample Guarantees and Connection to RegularizationHyuk Park, Grani A. Hanasusanto, Yingying LiICLR 2025
Builds on12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Improving Robustness using Generated DataSven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg et al.NeurIPS 2021 · 384 citations
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
Related papers
- A Unified Wasserstein Distributional Robustness Framework for Adversarial TrainingAnh Tuan Bui, Trung Le, Quan Hung Tran, He Zhao et al.ICLR 2022 · 54 citations
- Stronger and Faster Wasserstein Adversarial AttacksKaiwen Wu, Allen Houze Wang, Yaoliang YuICML 2020 · 42 citations
- Distributed Distributionally Robust Optimization with Non-Convex ObjectivesYang Jiao, Kai Yang, Dongjin SongNeurIPS 2022 · 21 citations
- Modeling the Second Player in Distributionally Robust OptimizationPaul Michel, Tatsunori Hashimoto, Graham NeubigICLR 2021 · 39 citations
- Provable Robust Overfitting Mitigation in Wasserstein Distributionally Robust OptimizationShuang Liu, Yihan Wang, Yifan Zhu, Yibo Miao et al.ICLR 2025
