Lune

NeurIPS2023Top-tier venue

Revisiting Adversarial Training for ImageNet: Architectures, Training and Generalization across Threat Models

Naman Deep Singh, Francesco Croce, Matthias Hein

2023Year
119Citations
40Top-tier citations

Abstract

While adversarial training has been extensively studied for ResNet architectures and low resolution datasets like CIFAR, much less is known for ImageNet. Given the recent debate about whether transformers are more robust than convnets, we revisit adversarial training on ImageNet comparing ViTs and ConvNeXts. Extensive experiments show that minor changes in architecture, most notably replacing PatchStem with ConvStem, and training scheme have a significant impact on the achieved robustness. These changes not only increase robustness in the seen ℓ∞\ell_\infty-threat model, but even more so improve generalization to unseen ℓ1/ℓ2\ell_1/\ell_2-attacks. Our modified ConvNeXt, ConvNeXt + ConvStem, yields the most robust ℓ∞\ell_\infty-models across different ranges of model parameters and FLOPs, while our ViT + ConvStem yields the best generalization to unseen threat models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ae571e49-4030-4aa7-b467-0ef91f09c367

Cited by top-tier papers40

Ask how each one uses it

Builds on35

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines