AGAIN: Adversarial Training with Attribution Span Enlargement and Hybrid Feature Fusion
Shenglin Yin, Kelu Yao, Sheng Shi, Yangzhou Du, Zhen Xiao
Abstract
The deep neural networks (DNNs) trained by adversarial training (AT) usually suffered from significant robust generalization gap, i.e., DNNs achieve high training robustness but low test robustness. In this paper, we propose a generic method to boost the robust generalization of AT methods from the novel perspective of attribution span. To this end, compared with standard DNNs, we discover that the generalization gap of adversarially trained DNNs is caused by the smaller attribution span on the input image. In other words, adversarially trained DNNs tend to focus on specific visual concepts on training images, causing its limitation on test robustness. In this way, to enhance the robustness, we propose an effective method to enlarge the learned attribution span. Besides, we use hybrid feature statistics for feature fusion to enrich the diversity of features. Extensive experiments show that our method can effectively improves robustness of adversarially trained DNNs, outperforming previous SOTA methods. Furthermore, we provide a theoretical analysis of our method to prove its effectiveness. * Zhen Xiao and Kelu Yao are the corresponding authors. (a) (b) (c) (d) (e)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26d0d208-77a0-46d7-ad3b-a2002ae3e279Cited by top-tier papers2
- Towards Adversarial Robustness via Debiased High-Confidence Logit AlignmentKejia Zhang, Juanjuan Weng, Shaozi Li, Zhiming LuoICCV 2025
- Nasty Adversarial Training: A Probability Sparsity Perspective for Robustness EnhancementYuhang Zhou, Zhongyun Hua, Zhaoquan Gu, Keke Tang et al.ICLR 2026
Builds on10
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Attacks Which Do Not Kill Training Make Adversarial Learning StrongerJingfeng Zhang, Xilie Xu, Bo Han, Gang Niu et al.ICML 2020 · 452 citations
Related papers
- Identifying and Understanding Cross-Class Features in Adversarial TrainingZeming Wei, Steven Y. Guo, Yisen WangICML 2025
- Consistency Regularization for Adversarial RobustnessJihoon Tack, Sihyun Yu, Jongheon Jeong, Minseon Kim et al.AAAI 2022 · 75 citations
- Enhanced Regularizers for Attributional RobustnessAnindya Sarkar, Anirban Sarkar, Vineeth N. BalasubramanianAAAI 2021 · 18 citations
- Balancing Generalization and Robustness in Adversarial Training via Steering through Clean and Adversarial Gradient DirectionsHaoyu Tong, Xiaoyu Zhang, Yulin Jin, Jian Lou et al.ACM MM 2024 · 2 citations
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra et al.ICLR 2022 · 54 citations
