Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness Tradeoff
Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai, Naoki Makishima, Atsushi Ando, Ryo Masumura
Abstract
This paper addresses the tradeoff between standard accuracy on clean examples and robustness against adversarial examples in deep neural networks (DNNs). Although adversarial training (AT) improves robustness, it degrades the standard accuracy, thus yielding the tradeoff. To mitigate this tradeoff, we propose a novel AT method called ARREST, which comprises three components: (i) adversarial finetuning (AFT), (ii) representation-guided knowledge distillation (RGKD), and (iii) noisy replay (NR). AFT trains a DNN on adversarial examples by initializing its parameters with a DNN that is standardly pretrained on clean examples. RGKD and NR respectively entail a regularization term and an algorithm to preserve latent representations of clean examples during AFT. RGKD penalizes the distance between the representations of the standardly pretrained and AFT DNNs. NR switches input adversarial examples to nonadversarial ones when the representation changes significantly during AFT. By combining these components, ARREST achieves both high standard accuracy and robustness. Experimental results demonstrate that ARREST mitigates the tradeoff more effectively than previous AT-based methods do.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70476c06-2b11-4e68-9ed6-12ef7e455217Cited by top-tier papers4
- Mitigating Feature Gap for Adversarial Robustness by Feature DisentanglementNuoyan Zhou, Dawei Zhou, Decheng Liu, Nannan Wang et al.AAAI 2025 · 3 citations
- Adversarially Pretrained Transformers May Be Universally Robust In-Context LearnersSoichiro Kumano, Hiroshi Kera, Toshihiko YamasakiICLR 2026 · 2 citations
- Sustainable Self-evolution Adversarial TrainingWenxuan Wang, Chenglei Wang, Huihui Qi, Menghao Ye et al.ACM MM 2024 · 2 citations
- Rethinking Invariance Regularization in Adversarial Training to Improve Robustness-Accuracy Trade-offFuta Kai Waseda, Ching-Chun Chang, Isao EchizenICLR 2025
Builds on21
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- Balancing Generalization and Robustness in Adversarial Training via Steering through Clean and Adversarial Gradient DirectionsHaoyu Tong, Xiaoyu Zhang, Yulin Jin, Jian Lou et al.ACM MM 2024 · 2 citations
- AGKD-BML: Defense Against Adversarial Attack by Attention Guided Knowledge Distillation and Bi-directional Metric LearningHong Wang, Yuefan Deng, Shinjae Yoo, Haibin Ling et al.ICCV 2021 · 21 citations
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 1 citation
- Exploring The Forgetting in Adversarial Training: A Novel Method for Enhancing RobustnessXianglu Wang, Hu DingICLR 2025
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 38 citations
