Adversarial Training for Probabilistic Robustness
Yi Zhang, Yuhang Chen, Zhen Chen, Wenjie Ruan, Xiaowei Huang, Siddartha Khastgir, Xingyu Zhao
Abstract
Deep learning (DL) has shown transformative potential across industries, yet its sensitivity to adversarial examples (AEs) limits its reliability and broader deployment. Research on DL robustness has developed various techniques, with adversarial training (AT) established as a leading approach to counter AEs. Traditional AT focuses on worst-case robustness (WCR), but recent work has introduced probabilistic robustness (PR), which evaluates the likelihood of AEs within a local perturbation range, providing an overall assessment of the model's robustness and acknowledging residual risks that are more practical to manage. However, existing AT methods are fundamentally designed to improve WCR, and no dedicated methods currently target PR. To bridge this gap, we formulate a new min-max optimization as the theoretical foundation for PR-focused AT, and introduce an AT-PR training scheme with numerical algorithms to solve the new optimization problem. Our experiments, based on 70 DL models trained on common datasets and diverse architectures, demonstrate that: i) AT-PR achieves higher improvements in PR than AT-WCR methods; ii) it shows more consistent effectiveness across varying local inputs; iii) it exhibits a reduced trade-off in model's generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1403f8ef-54ad-437c-bb57-03eafa68f2beBuilds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
Related papers
- Probabilistically Robust Learning: Balancing Average and Worst-case PerformanceAlexander Robey, Luiz F. O. Chamon, George J. Pappas, Hamed HassaniICML 2022 · 50 citations
- A Unified Wasserstein Distributional Robustness Framework for Adversarial TrainingAnh Tuan Bui, Trung Le, Quan Hung Tran, He Zhao et al.ICLR 2022 · 54 citations
- Non-Parametric Probabilistic Robustness: A Conservative Risk Estimator under Unknown Perturbation DistributionsZheng Wang, Yi Zhang, Siddartha Khastgir, carsten maple et al.ICML 2026
- Toward Improving the Robustness of Deep Learning Models via Model TransformationYingyi Zhang, Zan Wang, Jiajun Jiang, Hanmo You et al.ASE 2022 · 7 citations
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 15 citations
