Why adversarial training can hurt robust accuracy
Jacob Clarysse, Julia Hörrmann, Fanny Yang
摘要
Machine learning classifiers with high test accuracy often perform poorly under adversarial attacks. It is commonly believed that adversarial training alleviates this issue. In this paper, we demonstrate that, surprisingly, the opposite may be true -- Even though adversarial training helps when enough data is available, it may hurt robust generalization in the small sample size regime. We first prove this phenomenon for a high-dimensional linear classification setting with noiseless observations. Our proof provides explanatory insights that may also transfer to feature learning models. Further, we observe in experiments on standard image datasets that the same behavior occurs for perceptible attacks that effectively reduce class information such as mask attacks and object corruptions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Phase-aware Adversarial Defense for Improving Adversarial RobustnessDawei Zhou, Nannan Wang, Heng Yang, Xinbo Gao 等ICML 2023 · 被引用 14 次
- Theoretical Analysis of Robust Overfitting for Wide DNNs: An NTK ApproachShaopeng Fu, Di WangICLR 2024 · 被引用 9 次
- On the existence of consistent adversarial attacks in high-dimensional linear classificationMatteo Vilucchio, Lenka Zdeborova, Bruno LoureiroICML 2026 · 被引用 1 次
- Improving the Robustness of Transformer-based Large Language Models with Dynamic AttentionLujia Shen, Yuwen Pu, Shouling Ji, Changjiang Li 等NDSS 2024
- Phase and Amplitude-aware Prompting for Enhancing Adversarial RobustnessYibo Xu, Dawei Zhou, Decheng Liu, Nannan WangICML 2025
它引用的顶会 Paper18
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann 等NeurIPS 2020 · 被引用 688 次
- Adversarial Domain Adaptation with Domain MixupMinghao Xu, Jian Zhang, Bingbing Ni, Teng Li 等AAAI 2020 · 被引用 499 次
相关 Paper
- More Data Can Expand The Generalization Gap Between Adversarially Robust and Standard ModelsLin Chen, Yifei Min, Mingrui Zhang, Amin KarbasiICML 2020 · 被引用 66 次
- Phase Transition from Clean Training to Adversarial TrainingYue Xing, Qifan Song, Guang ChengNeurIPS 2022 · 被引用 4 次
- Defending Against Universal Perturbations With Shared Adversarial TrainingChaithanya Kumar Mummadi, Thomas Brox, Jan Hendrik MetzenICCV 2019 · 被引用 61 次
- Precise Accuracy / Robustness Tradeoffs in Regression: Case of General NormsElvis Dohmatob, Meyer ScetbonICML 2024 · 被引用 3 次
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang 等ICLR 2024
