Pyramid Adversarial Training Improves ViT Performance
Charles Herrmann, Kyle Sargent, Lu Jiang, Ramin Zabih, Huiwen Chang, Ce Liu, Dilip Krishnan, Deqing Sun
摘要
Aggressive data augmentation is a key component of the strong generalization capabilities of Vision Transformer (ViT). One such data augmentation technique is adversarial training (AT); however, many prior works [28, 45] have shown that this often results in poor clean accuracy. In this work, we present pyramid adversarial training (Pyra-midAT), a simple and effective technique to improve ViT's overall performance. We pair it with a "matched" Dropout and stochastic depth regularization, which adopts the same Dropout and stochastic depth configuration for the clean and adversarial samples. Similar to the improvements on CNNs by AdvProp [61] (not directly applicable to ViT), our pyramid adversarial training breaks the trade-off between in-distribution accuracy and out-of-distribution robustness for ViT and related architectures. It leads to 1.82% absolute improvement on ImageNet clean accuracy for the ViT-B model when trained only on ImageNet-1K data, while simultaneously boosting performance on 7 ImageNet robustness metrics, by absolute numbers ranging from 1.76% to 15.68%. We set a new state-of-the-art for ImageNet-C (41.42 mCE), ImageNet-R (53.92%), and ImageNet-Sketch (41.04%) without extra data, using only the ViT-B/16 backbone and our pyramid adversarial training. Our code is publicly available at pyramidat.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Zero-Shot Learning by Harnessing Adversarial SamplesZhi Chen, Peng-Fei Zhang, Jingjing Li, Sen Wang 等ACM MM 2023 · 被引用 29 次
- LPT: Long-tailed Prompt Tuning for Image ClassificationBowen Dong, Pan Zhou, Shuicheng Yan, Wangmeng ZuoICLR 2023 · 被引用 19 次
- Distilling Out-of-Distribution Robustness from Vision-Language Foundation ModelsAndy Zhou, Jindong Wang, Yu-Xiong Wang, Haohan WangNeurIPS 2023 · 被引用 14 次
- High-dimensional (Group) Adversarial Training in Linear RegressionYiling Xie, Xiaoming HuoNeurIPS 2024 · 被引用 8 次
- Understanding and Defending Patched-based Adversarial Attacks for Vision TransformerLiang Liu, Yanan Guo, Youtao Zhang, Jun YangICML 2023 · 被引用 7 次
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
相关 Paper
- Revisiting adapters with adversarial trainingSylvestre-Alvise Rebuffi, Francesco Croce, Sven GowalICLR 2023
- Towards Robust Vision TransformerXiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li 等CVPR 2022 · 被引用 185 次
- Scale-space Tokenization for Improving the Robustness of Vision TransformersLei Xu, Rei Kawakami, Nakamasa InoueACM MM 2023 · 被引用 1 次
- Vision Transformers Are Robust LearnersSayak Paul, Pin-Yu ChenAAAI 2022 · 被引用 372 次
- Improving robustness to corruptions with multiplicative weight perturbationsTrung Q. Trinh, Markus Heinonen, Luigi Acerbi, Samuel KaskiNeurIPS 2024 · 被引用 8 次
