Adversarial Distillation Based on Slack Matching and Attribution Region Alignment
Shenglin Yin, Zhen Xiao, Mingxuan Song, Jieyi Long
Abstract
Adversarial distillation (AD) is a highly effective method for enhancing the robustness of small models. Contrary to expectations, a high-performing teacher model does not always result in a more robust student model. This is due to two main reasons. First, when there are significant differences in predictions between the teacher model and the student model, exact matching of predicted values using KL divergence interferes with training, leading to poor performance of existing methods. Second, matching solely based on the output prevents the student model from fully understanding the behavior of the teacher model. To address these challenges, this paper proposes a novel AD method named SmaraAD. During the training process, we facilitate the student model in better understanding the teacher model's behavior by aligning the attribution region that the student model focuses on with that of the teacher model. Concurrently, we relax the condition of exact matching in KL divergence and replace it with a more flexible matching criterion, thereby enhancing the model's robustness. Extensive experiments substantiate the effectiveness of our method in improving the robustness of small models, out-performing previous SOTA methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28134f3a-df1f-4b2e-a860-5b26e13bbf65Cited by top-tier papers4
- Indirect Gradient Matching for Adversarial Robust DistillationHongsin Lee, Seungju Cho, Changick KimICLR 2025
- Nasty Adversarial Training: A Probability Sparsity Perspective for Robustness EnhancementYuhang Zhou, Zhongyun Hua, Zhaoquan Gu, Keke Tang et al.ICLR 2026
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
- Rethinking the Dark Knowledge and Kullback-Leibler Divergence Loss in Knowledge Distillation Under Capacity MismatchingYingchao Wang, Wenqi Niu, Xingshan Yao, Li You et al.AAAI 2026
Builds on18
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
Related papers
- Adversarially Robust DistillationMicah Goldblum, Liam Fowl, Soheil Feizi, Tom GoldsteinAAAI 2020 · 258 citations
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student BetterBojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang JiangICCV 2021 · 136 citations
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 8 citations
- Robust Distillation via Untargeted and Targeted Intermediate Adversarial SamplesJunhao Dong, Piotr Koniusz, Junxi Chen, Z. Jane Wang et al.CVPR 2024
- Advantage-Guided Distillation for Preference Alignment in Small Language ModelsShiping Gao, Fanqi Wan, Jiajian Guo, Xiaojun Quan et al.ICLR 2025
