Identifying and Understanding Cross-Class Features in Adversarial Training
Zeming Wei, Steven Y. Guo, Yisen Wang
Abstract
Adversarial training (AT) has been considered one of the most effective methods for making deep neural networks robust against adversarial attacks, while the training mechanisms and dynamics of AT remain open research problems. In this paper, we present a novel perspective on studying AT through the lens of class-wise feature attribution. Specifically, we identify the impact of a key family of features on AT that are shared by multiple classes, which we call crossclass features. These features are typically useful for robust classification, which we offer theoretical evidence to illustrate through a synthetic data model. Through systematic studies across multiple model architectures and settings, we find that during the initial stage of AT, the model tends to learn more cross-class features until the best robustness checkpoint. As AT further squeezes the training robust loss and causes robust overfitting, the model tends to make decisions based on more class-specific features. Based on these discoveries, we further provide a unified view of two existing properties of AT, including the advantage of soft-label training and robust overfitting. Overall, these insights refine the current understanding of AT mechanisms and provide new perspectives on studying them. Our code is available at https://github.com/ PKU-ML/Cross-Class-Features-AT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91b89575-08a0-4e49-9885-d21474359a69Builds on34
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- CFA: Class-Wise Calibrated Fair Adversarial TrainingZeming Wei, Yifei Wang, Yiwen Guo, Yisen WangCVPR 2023
- Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured DataBinghui Li, Yuanzhi LiICLR 2025
- Relating Adversarially Robust Generalization to Flat MinimaDavid Stutz, Matthias Hein, Bernt SchieleICCV 2021 · 80 citations
- Advancing Example Exploitation Can Alleviate Critical Challenges in Adversarial TrainingYao Ge, Yun Li, Keji Han, Junyi Zhu et al.ICCV 2023 · 6 citations
- AGAIN: Adversarial Training with Attribution Span Enlargement and Hybrid Feature FusionShenglin Yin, Kelu Yao, Sheng Shi, Yangzhou Du et al.CVPR 2023
