AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-directional Metric Learning
Hong Wang, Yuefan Deng, Shinjae Yoo, Haibin Ling, Yuewei Lin
Abstract
While deep neural networks have shown impressive performance in many tasks, they are fragile to carefully de-signed adversarial attacks. We propose a novel adversarial training-based model by Attention Guided Knowledge Distillation and Bi-directional Metric Learning (AGKD-BML). The attention knowledge is obtained from a weight-fixed model trained on a clean dataset, referred to as a teacher model, and transferred to a model that is under training on adversarial examples (AEs), referred to as a student model. In this way, the student model is able to focus on the correct region, as well as correcting the intermediate features corrupted by AEs to eventually improve the model accuracy. Moreover, to efficiently regularize the representation in feature space, we propose a bidirectional metric learning. Specifically, given a clean image, it is first attacked to its most confusing class to get the forward AE. A clean image in the most confusing class is then randomly picked and attacked back to the original class to get the backward AE. A triplet loss is then used to shorten the representation distance between original image and its AE, while enlarge that between the forward and backward AEs. We conduct extensive adversarial robustness experiments on two widely used datasets with different attacks. Our proposed AGKD-BML model consistently outperforms the state-of-the-art approaches. The code of AGKD-BML will be available at: https://github.com/hongw579/AGKD-BML.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6df7a2a4-022d-4b52-8b39-afe5c31c9e2fCited by top-tier papers2
- Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness TradeoffSatoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai et al.ICCV 2023 · 8 citations
- Defense without Forgetting: Continual Adversarial Defense with Anisotropic & Isotropic Pseudo ReplayYuhang Zhou, Zhongyun HuaCVPR 2024 · 3 citations
Builds on13
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
Related papers
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
- Improving Adversarial Robust Fairness via Anti-Bias Soft Label DistillationShiji Zhao, Ranjie Duan, Xizhe Wang, Xingxing WeiNeurIPS 2024 · 12 citations
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student BetterBojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang JiangICCV 2021 · 136 citations
- Two-Stage Adversarial Training for Deep Hashing via Representation DistillationFei Zhu, Huashan Chen, Wanqian Zhang, Lin Wang et al.SIGIR 2025 · 2 citations
- Improving Adversarial Robustness via Information Bottleneck DistillationHuafeng Kuang, Hong Liu, Yongjian Wu, Shin'ichi Satoh et al.NeurIPS 2023 · 27 citations
