Adversarial Training and Provable Robustness: A Tale of Two Objectives
Jiameng Fan, Wenchao Li
Abstract
We propose a principled framework that combines adversarial training and provable robustness verification for training certifiably robust neural networks. We formulate the training problem as a joint optimization problem with both empirical and provable robustness objectives and develop a novel gradient-descent technique that can eliminate bias in stochastic multi-gradients. We perform both theoretical analysis on the convergence of the proposed technique and experimental comparison with state-of-the-arts. Results on MNIST and CIFAR-10 show that our method can consistently match or outperform prior approaches for provable l∞ robustness. Notably, we achieve 6.60% verified test error on MNIST at ε = 0.3, and 66.57% on CIFAR-10 with ε = 8/255.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4929998-cdd6-48a5-9401-e98a86df416eCited by top-tier papers4
- Expressive Losses for Verified Robustness via Convex CombinationsAlessandro De Palma, Rudy Bunel, Krishnamurthy (Dj) Dvijotham, M. Pawan Kumar et al.ICLR 2024 · 27 citations
- REGLO: Provable Neural Network Repair for Global Robustness PropertiesFeisi Fu, Zhilu Wang, Weichao Zhou, Yixuan Wang et al.AAAI 2024 · 11 citations
- Support is All You Need for Certified VAE TrainingChangming Xu, Debangshu Banerjee, Deepak Vasisht, Gagandeep SinghICLR 2025
- Boosting Verified Training for Robust Image Classifications via AbstractionZhaodi Zhang, Zhiyi Xue, Yang Chen, Si Liu et al.CVPR 2023
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Adversarial Training and Provable Defenses: Bridging the GapMislav Balunovic, Martin T. VechevICLR 2020 · 186 citations
Related papers
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel et al.ICCV 2019 · 196 citations
- Connecting Certified and Adversarial TrainingYuhao Mao, Mark Niklas Müller, Marc Fischer, Martin T. VechevNeurIPS 2023 · 14 citations
- Rethinking Evaluation Paradigms in IBP-based Certified TrainingKonstantin Kaulen, Hadar Shavit, Holger HoosICML 2026
- Globally-Robust Neural NetworksKlas Leino, Zifan Wang, Matt FredriksonICML 2021 · 150 citations
