What It Thinks Is Important Is Important: Robustness Transfers Through Input Gradients
Alvin Chan, Yi Tay, Yew-Soon Ong
Abstract
Adversarial perturbations are imperceptible changes to input pixels that can change the prediction of deep learning models. Learned weights of models robust to such perturbations are previously found to be transferable across different tasks but this applies only if the model architecture for the source and target tasks is the same. Input gradients characterize how small changes at each input pixel affect the model output. Using only natural images, we show here that training a student model's input gradients to match those of a robust teacher model can gain robustness close to a strong baseline that is robustly trained from scratch. Through experiments in MNIST, CIFAR-10, CIFAR-100 and Tiny-ImageNet, we show that our proposed method, input gradient adversarial matching, can transfer robustness across different tasks and even across different model architectures. This demonstrates that directly targeting the semantics of input gradients is a feasible way towards adversarial robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06c1aa8e-25c2-45d2-85f8-122f926832e6Cited by top-tier papers17
- Adversarial Robustness for Unsupervised Domain AdaptationMuhammad Awais, Fengwei Zhou, Hang Xu, Lanqing Hong et al.ICCV 2021 · 46 citations
- Federated Robustness Propagation: Sharing Adversarial Robustness in Heterogeneous Federated LearningJunyuan Hong, Haotao Wang, Zhangyang Wang, Jiayu ZhouAAAI 2023 · 29 citations
- Removing Undesirable Feature Contributions Using Out-of-Distribution DataSaehyung Lee, Changhwa Park, Hyungyu Lee, Jihun Yi et al.ICLR 2021 · 26 citations
- Does Robustness on ImageNet Transfer to Downstream Tasks?Yutaro Yamada, Mayu OtaniCVPR 2022 · 23 citations
- MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel MapsMuhammad Awais, Fengwei Zhou, Chuanlong Xie, Jiawei Li et al.NeurIPS 2021 · 22 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
- COCO-GAN: Generation by Parts via Conditional CoordinatingChieh Hubert Lin, Chia-Che Chang, Yu-Sheng Chen, Da-Cheng Juan et al.ICCV 2019 · 147 citations
- Adversarially robust transfer learningAli Shafahi, Parsa Saadatpanah, Chen Zhu, Amin Ghiasi et al.ICLR 2020 · 130 citations
Related papers
- Transferring Adversarial Robustness Through Robust Representation MatchingPratik Vaishnavi, Kevin Eykholt, Amir RahmatiUSENIX Security 2022
- Indirect Gradient Matching for Adversarial Robust DistillationHongsin Lee, Seungju Cho, Changick KimICLR 2025
- A Little Robustness Goes a Long Way: Leveraging Robust Features for Targeted Transfer AttacksJacob M. Springer, Melanie Mitchell, Garrett T. KenyonNeurIPS 2021 · 54 citations
- Robust Weight Signatures: Gaining Robustness as Easy as Patching Weights?Ruisi Cai, Zhenyu Zhang, Zhangyang WangICML 2023 · 16 citations
- Do Perceptually Aligned Gradients Imply Robustness?Roy Ganz, Bahjat Kawar, Michael EladICML 2023 · 18 citations
