Improving Robustness using Generated Data
Sven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg, Dan Andrei Calian, Timothy A. Mann
Abstract
Recent work argues that robust training requires substantially larger datasets than those required for standard classification. On CIFAR-10 and CIFAR-100, this translates into a sizable robust-accuracy gap between models trained solely on data from the original training set and those trained with additional data extracted from the"80 Million Tiny Images"dataset (TI-80M). In this paper, we explore how generative models trained solely on the original training set can be leveraged to artificially increase the size of the original training set and improve adversarial robustness to norm-bounded perturbations. We identify the sufficient conditions under which incorporating additional generated data can improve robustness, and demonstrate that it is possible to significantly reduce the robust-accuracy gap to models trained with additional real data. Surprisingly, we even show that even the addition of non-realistic random data (generated by Gaussian sampling) can improve robustness. We evaluate our approach on CIFAR-10, CIFAR-100, SVHN and TinyImageNet against and norm-bounded perturbations of size and , respectively. We show large absolute improvements in robust accuracy compared to previous state-of-the-art methods. Against norm-bounded perturbations of size , our models achieve 66.10% and 33.49% robust accuracy on CIFAR-10 and CIFAR-100, respectively (improving upon the state-of-the-art by +8.96% and +3.29%). Against norm-bounded perturbations of size , our model achieves 78.31% on CIFAR-10 (+3.81%). These results beat most prior works that use external data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d97e4a2-17d5-4a63-b4ba-6f067ec243e3Cited by top-tier papers120
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
- Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionTianyu Pang, Min Lin, Xiao Yang, Jun Zhu et al.ICML 2022 · 168 citations
- Decoupled Kullback-Leibler Divergence LossJiequan Cui, Zhuotao Tian, Zhisheng Zhong, Xiaojuan Qi et al.NeurIPS 2024 · 119 citations
- Revisiting Adversarial Training for ImageNet: Architectures, Training and Generalization across Threat ModelsNaman Deep Singh, Francesco Croce, Matthias HeinNeurIPS 2023 · 119 citations
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
Related papers
- Data Augmentation Can Improve RobustnessSylvestre-Alvise Rebuffi, Sven Gowal, Dan Andrei Calian, Florian Stimberg et al.NeurIPS 2021 · 427 citations
- More Data Can Expand The Generalization Gap Between Adversarially Robust and Standard ModelsLin Chen, Yifei Min, Mingrui Zhang, Amin KarbasiICML 2020 · 66 citations
- Adapting to Evolving Adversaries with Regularized Continual Robust TrainingSihui Dai, Christian Cianfarani, Vikash Sehwag, Prateek Mittal et al.ICML 2025
- Robust Learning Meets Generative Models: Can Proxy Distributions Improve Adversarial Robustness?Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina, Sihui Dai et al.ICLR 2022 · 150 citations
- On the Scalability of Certified Adversarial Robustness with Generated DataThomas Altstidl, David Dobre, Arthur Kosmala, Bjoern M. Eskofier et al.NeurIPS 2024 · 10 citations
