Reducing Excessive Margin to Achieve a Better Accuracy vs. Robustness Trade-off
Rahul Rade, Seyed-Mohsen Moosavi-Dezfooli
Abstract
While adversarial training has become the de facto approach for training robust classifiers, it leads to a drop in accuracy. This has led to prior works postulating that accuracy is inherently at odds with robustness. Yet, the phenomenon remains inexplicable. In this paper, we closely examine the changes induced in the decision boundary of a deep network during adversarial training. We find that adversarial training leads to unwarranted increase in the margin along certain adversarial directions, thereby hurting accuracy. Motivated by this observation, we present a novel algorithm, called Helper-based Adversarial Training (HAT), to reduce this effect by incorporating additional wrongly labelled examples during training. Our proposed method provides a notable improvement in accuracy without compromising robustness. It achieves a better trade-off between accuracy and robustness in comparison to existing defenses. Code is available at https://github.com/imrahulr/hat.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 228eb77c-34b6-4863-8a40-93f29f34ac24Cited by top-tier papers32
- Revisiting Adversarial Training for ImageNet: Architectures, Training and Generalization across Threat ModelsNaman Deep Singh, Francesco Croce, Matthias HeinNeurIPS 2023 · 119 citations
- Robust Evaluation of Diffusion-Based Adversarial PurificationMinjong Lee, Dongwoo KimICCV 2023 · 96 citations
- Efficient and Effective Augmentation Strategy for Adversarial TrainingSravanti Addepalli, Samyak Jain, Venkatesh Babu R.NeurIPS 2022 · 77 citations
- Enhance the Visual Representation via Discrete Adversarial TrainingXiaofeng Mao, Yuefeng Chen, Ranjie Duan, Yao Zhu et al.NeurIPS 2022 · 48 citations
- Your Out-of-Distribution Detection Method is Not Robust!Mohammad Azizmalayeri, Arshia Soltani Moakhar, Arman Zarei, Reihaneh Zohrabi et al.NeurIPS 2022 · 29 citations
Related papers
- Splitting the Difference on Adversarial TrainingMatan Levi, Aryeh KontorovichUSENIX Security 2024 · 9 citations
- Geometry-aware Instance-reweighted Adversarial TrainingJingfeng Zhang, Jianing Zhu, Gang Niu, Bo Han et al.ICLR 2021 · 316 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingYanyun Wang, Li LiuICCV 2025 · 1 citation
- Exploring and Exploiting Decision Boundary Dynamics for Adversarial RobustnessYuancheng Xu, Yanchao Sun, Micah Goldblum, Tom Goldstein et al.ICLR 2023 · 10 citations
