Wide Two-Layer Networks can Learn from Adversarial Perturbations
Soichiro Kumano, Hiroshi Kera, Toshihiko Yamasaki
Abstract
Adversarial examples have raised several open questions, such as why they can deceive classifiers and transfer between different models. A prevailing hypothesis to explain these phenomena suggests that adversarial perturbations appear as random noise but contain class-specific features. This hypothesis is supported by the success of perturbation learning, where classifiers trained solely on adversarial examples and the corresponding incorrect labels generalize well to correctly labeled test data. Although this hypothesis and perturbation learning are effective in explaining intriguing properties of adversarial examples, their solid theoretical foundation is limited. In this study, we theoretically explain the counterintuitive success of perturbation learning. We assume wide two-layer networks and the results hold for any data distribution. We prove that adversarial perturbations contain sufficient class-specific features for networks to generalize from them. Moreover, the predictions of classifiers trained on mislabeled adversarial examples coincide with those of classifiers trained on correctly labeled clean samples. The code is available at https://github.com/s-kumano/perturbation-learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 129ee9cc-3504-4cc2-9fa3-3dcfa5236f2bBuilds on14
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor et al.NeurIPS 2020 · 506 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov et al.NeurIPS 2020 · 336 citations
- Understanding and Mitigating the Tradeoff between Robustness and AccuracyAditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi et al.ICML 2020 · 252 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
Related papers
- Theoretical Understanding of Learning from Adversarial PerturbationsSoichiro Kumano, Hiroshi Kera, Toshihiko YamasakiICLR 2024 · 4 citations
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 83 citations
- Towards Transferable Targeted Adversarial ExamplesZhibo Wang, Hongshan Yang, Yunhe Feng, Peng Sun et al.CVPR 2023
- Modeling Adversarial Noise for Adversarial TrainingDawei Zhou, Nannan Wang, Bo Han, Tongliang LiuICML 2022 · 20 citations
- Introducing Competition to Boost the Transferability of Targeted Adversarial Examples Through Clean Feature MixupJunyoung Byun, Myung-Joon Kwon, Seungju Cho, Yoonji Kim et al.CVPR 2023
