Transferring Adversarial Robustness Through Robust Representation Matching
Pratik Vaishnavi, Kevin Eykholt, Amir Rahmati
摘要
With the widespread use of machine learning, concerns over its security and reliability have become prevalent. As such, many have developed defenses to harden neural networks against adversarial examples, imperceptibly perturbed inputs that are reliably misclassified. Adversarial training in which adversarial examples are generated and used during training is one of the few known defenses able to reliably withstand such attacks against neural networks. However, adversarial training imposes a significant training overhead and scales poorly with model complexity and input dimension. In this paper, we propose Robust Representation Matching (RRM), a low-cost method to transfer the robustness of an adversarially trained model to a new model being trained for the same task irrespective of architectural differences. Inspired by student-teacher learning, our method introduces a novel training loss that encourages the student to learn the teacher's robust representations. Compared to prior works, RRM is superior with respect to both model performance and adversarial training time. On CIFAR-10, RRM trains a robust model faster than the state-of-the-art. Furthermore, RRM remains effective on higher-dimensional datasets. On Restricted-ImageNet, RRM trains a ResNet50 model faster than standard adversarial training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- Understanding Zero-shot Adversarial Robustness for Large-Scale ModelsChengzhi Mao, Scott Geng, Junfeng Yang, Xin Wang 等ICLR 2023 · 被引用 10 次
- Ciard: Cyclic Iterative Adversarial Robustness DistillationLiming Lu, Shuchao Pang, Xu Zheng, Xiang Gu 等ICCV 2025 · 被引用 1 次
- Toward Understanding Adversarial Distillation: Why Robust Teachers FailHongsin Lee, Hye Won ChungICML 2026
- Boosting Accuracy and Robustness of Student Models via Adaptive Adversarial DistillationBo Huang, Mingyang Chen, Yi Wang, Junda Lu 等CVPR 2023
它引用的顶会 Paper7
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
相关 Paper
- Accelerating Certified Robustness Training via Knowledge TransferPratik Vaishnavi, Kevin Eykholt, Amir RahmatiNeurIPS 2022 · 被引用 8 次
- What It Thinks Is Important Is Important: Robustness Transfers Through Input GradientsAlvin Chan, Yi Tay, Yew-Soon OngCVPR 2020
- MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel MapsMuhammad Awais, Fengwei Zhou, Chuanlong Xie, Jiawei Li 等NeurIPS 2021 · 被引用 22 次
- Efficient Adversarial Training With Transferable Adversarial ExamplesHaizhong Zheng, Ziqi Zhang, Juncheng Gu, Honglak Lee 等CVPR 2020
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor 等NeurIPS 2020 · 被引用 506 次
