Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better
Bojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang Jiang
Abstract
Adversarial training is one effective approach for training robust deep neural networks against adversarial attacks. While being able to bring reliable robustness, adversarial training (AT) methods in general favor high capacity models, i.e., the larger the model the better the robustness. This tends to limit their effectiveness on small models, which are more preferable in scenarios where storage or computing resources are very limited (e.g., mobile devices). In this paper, we leverage the concept of knowledge distillation to improve the robustness of small models by distilling from adversarially trained large models. We first revisit several state-of-the-art AT methods from a distillation perspective and identify one common technique that can lead to improved robustness: the use of robust soft labels – predictions of a robust model. Following this observation, we propose a novel adversarial robustness distillation method called Robust Soft Label Adversarial Distillation (RSLAD) to train robust small student models. RSLAD fully exploits the robust soft labels produced by a robust (adversarially-trained) large teacher model to guide the student’s learning on both natural and adversarial examples in all loss terms. We empirically demonstrate the effectiveness of our RSLAD approach over existing adversarial training and distillation methods in improving the robustness of small models against state-of-the-art attacks including the AutoAttack. We also provide a set of understandings on our RSLAD and the importance of robust soft labels for adversarial robustness distillation. Code: https://github.com/zibojia/RSLAD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46fa2a39-0a4b-485b-87a6-e05bfbe0f608Cited by top-tier papers34
- Exploring Architectural Ingredients of Adversarially Robust Deep Neural NetworksHanxun Huang, Yisen Wang, Sarah M. Erfani, Quanquan Gu et al.NeurIPS 2021 · 124 citations
- Reliable Adversarial Distillation with Unreliable TeachersJianing Zhu, Jiangchao Yao, Bo Han, Jingfeng Zhang et al.ICLR 2022 · 92 citations
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra et al.ICLR 2022 · 54 citations
- Revisiting Adversarial Robustness Distillation from the Perspective of Robust FairnessXinli Yue, Ningping Mou, Qian Wang, Lingchen ZhaoNeurIPS 2023 · 28 citations
- Improving Adversarial Robustness via Information Bottleneck DistillationHuafeng Kuang, Hong Liu, Yongjian Wu, Shin'ichi Satoh et al.NeurIPS 2023 · 27 citations
Builds on17
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- Adversarially Robust DistillationMicah Goldblum, Liam Fowl, Soheil Feizi, Tom GoldsteinAAAI 2020 · 258 citations
- Improving Adversarial Robust Fairness via Anti-Bias Soft Label DistillationShiji Zhao, Ranjie Duan, Xizhe Wang, Xingxing WeiNeurIPS 2024 · 12 citations
- Indirect Gradient Matching for Adversarial Robust DistillationHongsin Lee, Seungju Cho, Changick KimICLR 2025
- Annealing Self-Distillation Rectification Improves Adversarial TrainingYu-Yu Wu, Hung-Jui Wang, Shang-Tse ChenICLR 2024 · 10 citations
- Boosting Accuracy and Robustness of Student Models via Adaptive Adversarial DistillationBo Huang, Mingyang Chen, Yi Wang, Junda Lu et al.CVPR 2023
