Toward Understanding Adversarial Distillation: Why Robust Teachers Fail
Hongsin Lee, Hye Won Chung
Abstract
Adversarial Distillation aims to enhance student robustness by guiding the student with a robust teacher's soft labels within the min-max adversarial training framework, yet its success is notoriously inconsistent: a more robust teacher often fails to improve, or even harms, the student's robust generalization. In this paper, we identify a key mechanism of this teacher dependency: the misalignment between the teacher's supervisory confidence and the student's representational limitations on a consistent subset of training data—the Robustly Unlearnable Set. We present a theoretical framework analyzing the feature learning dynamics of a two-layer neural network, demonstrating that this mismatch creates a dichotomy in distillation outcomes. We prove that when a teacher provides confident supervision on unlearnable samples, it compels the student to memorize spurious noise patterns that eventually overpower the learned robust signal, thereby driving robust overfitting. Conversely, a teacher that exhibits high uncertainty on these samples effectively suppresses noise memorization, allowing the student to rely solely on the learnable signal for robust generalization. We empirically validate our theory across both synthetic simulations and real-image classification datasets, confirming that robust overfitting is driven by the teacher's interaction with unlearnable samples. Finally, we demonstrate that a teacher's predictive entropy on unlearnable samples serves as a strong indicator of student robustness, validating our theoretical framework and offering a principled guideline for robust teacher selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on71
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
Related papers
- Reliable Adversarial Distillation with Unreliable TeachersJianing Zhu, Jiangchao Yao, Bo Han, Jingfeng Zhang et al.ICLR 2022 · 92 citations
- Soften to Defend: Towards Adversarial Robustness via Self-Guided Label RefinementZhuorong Li, Daiwei Yu, Lina Wei, Canghong Jin et al.CVPR 2024
- Boosting Accuracy and Robustness of Student Models via Adaptive Adversarial DistillationBo Huang, Mingyang Chen, Yi Wang, Junda Lu et al.CVPR 2023
- Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher KnowledgeFreya Behrens, Lenka ZdeborováICLR 2026 · 9 citations
- Adversarially Robust DistillationMicah Goldblum, Liam Fowl, Soheil Feizi, Tom GoldsteinAAAI 2020 · 258 citations
