What It Thinks Is Important Is Important: Robustness Transfers Through Input Gradients
Alvin Chan, Yi Tay, Yew-Soon Ong
摘要
Adversarial perturbations are imperceptible changes to input pixels that can change the prediction of deep learning models. Learned weights of models robust to such perturbations are previously found to be transferable across different tasks but this applies only if the model architecture for the source and target tasks is the same. Input gradients characterize how small changes at each input pixel affect the model output. Using only natural images, we show here that training a student model's input gradients to match those of a robust teacher model can gain robustness close to a strong baseline that is robustly trained from scratch. Through experiments in MNIST, CIFAR-10, CIFAR-100 and Tiny-ImageNet, we show that our proposed method, input gradient adversarial matching, can transfer robustness across different tasks and even across different model architectures. This demonstrates that directly targeting the semantics of input gradients is a feasible way towards adversarial robustness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Adversarial Robustness for Unsupervised Domain AdaptationMuhammad Awais, Fengwei Zhou, Hang Xu, Lanqing Hong 等ICCV 2021 · 被引用 46 次
- Federated Robustness Propagation: Sharing Adversarial Robustness in Heterogeneous Federated LearningJunyuan Hong, Haotao Wang, Zhangyang Wang, Jiayu ZhouAAAI 2023 · 被引用 29 次
- Removing Undesirable Feature Contributions Using Out-of-Distribution DataSaehyung Lee, Changhwa Park, Hyungyu Lee, Jihun Yi 等ICLR 2021 · 被引用 26 次
- Does Robustness on ImageNet Transfer to Downstream Tasks?Yutaro Yamada, Mayu OtaniCVPR 2022 · 被引用 23 次
- MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel MapsMuhammad Awais, Fengwei Zhou, Chuanlong Xie, Jiawei Li 等NeurIPS 2021 · 被引用 22 次
它引用的顶会 Paper4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 被引用 597 次
- COCO-GAN: Generation by Parts via Conditional CoordinatingChieh Hubert Lin, Chia-Che Chang, Yu-Sheng Chen, Da-Cheng Juan 等ICCV 2019 · 被引用 147 次
- Adversarially robust transfer learningAli Shafahi, Parsa Saadatpanah, Chen Zhu, Amin Ghiasi 等ICLR 2020 · 被引用 130 次
相关 Paper
- Transferring Adversarial Robustness Through Robust Representation MatchingPratik Vaishnavi, Kevin Eykholt, Amir RahmatiUSENIX Security 2022
- Indirect Gradient Matching for Adversarial Robust DistillationHongsin Lee, Seungju Cho, Changick KimICLR 2025
- A Little Robustness Goes a Long Way: Leveraging Robust Features for Targeted Transfer AttacksJacob M. Springer, Melanie Mitchell, Garrett T. KenyonNeurIPS 2021 · 被引用 54 次
- Robust Weight Signatures: Gaining Robustness as Easy as Patching Weights?Ruisi Cai, Zhenyu Zhang, Zhangyang WangICML 2023 · 被引用 16 次
- Do Perceptually Aligned Gradients Imply Robustness?Roy Ganz, Bahjat Kawar, Michael EladICML 2023 · 被引用 18 次
