Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness Tradeoff
Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando, Ryo Masumura
摘要
This paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff, we propose a novel AT method called ARREST, which comprises three components: (i) adversarial finetuning (AFT), (ii) representation-guided knowledge distillation (RGKD), and (iii) noisy replay (NR). AFT trains a DNN on adversarial examples by initializing its parameters with a DNN that is standardly pretrained on clean examples. RGKD and NR respectively entail a regularization term and an algorithm to preserve latent representations of clean examples during AFT. RGKD penalizes the distance between the representations of the standardly pretrained and AFT DNNs. NR switches input adversarial examples to nonadversarial ones when the representation changes significantly during AFT. By combining these components, ARREST achieves both high standard accuracy and robustness. Experimental results demonstrate that ARREST mitigates the tradeoff more effectively than previous AT-based methods do.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Mitigating Feature Gap for Adversarial Robustness by Feature DisentanglementNuoyan Zhou, Dawei Zhou, Decheng Liu, Nannan Wang 等AAAI 2025 · 被引用 3 次
- Adversarially Pretrained Transformers May Be Universally Robust In-Context LearnersSoichiro Kumano, Hiroshi Kera, Toshihiko YamasakiICLR 2026 · 被引用 2 次
- Sustainable Self-evolution Adversarial TrainingWenxuan Wang, Chenglei Wang, Huihui Qi, Menghao Ye 等ACM MM 2024 · 被引用 2 次
- Rethinking Invariance Regularization in Adversarial Training to Improve Robustness-Accuracy Trade-offFuta Kai Waseda, Ching-Chun Chang, Isao EchizenICLR 2025
它引用的顶会 Paper21
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
相关 Paper
- Balancing Generalization and Robustness in Adversarial Training via Steering through Clean and Adversarial Gradient DirectionsHaoyu Tong, Xiaoyu Zhang, Yulin Jin, Jian Lou 等ACM MM 2024 · 被引用 2 次
- AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-directional Metric LearningHong Wang, Yuefan Deng, Shinjae Yoo, Haibin Ling 等ICCV 2021 · 被引用 21 次
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 被引用 1 次
- Exploring The Forgetting in Adversarial Training: A Novel Method for Enhancing RobustnessXianglu Wang, Hu DingICLR 2025
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 被引用 38 次
