Improving Adversarial Robustness Requires Revisiting Misclassified Examples
Yisen Wang, Difan Zou, Jinfeng Yi, James Bailey, Xingjun Ma, Quanquan Gu
Abstract
Deep neural networks (DNNs) are vulnerable to adversarial examples crafted by imperceptible perturbations. A range of defense techniques have been proposed to improve DNN robustness to adversarial examples, among which adversarial training has been demonstrated to be the most effective. Adversarial training is often formulated as a min-max optimization problem, with the inner maximization for generating adversarial examples. However, there exists a simple, yet easily overlooked fact that adversarial examples are only defined on correctly classified (natural) examples, but inevitably, some (natural) examples will be misclassified during training. In this paper, we investigate the distinctive influence of misclassified and correctly classified examples on the final robustness of adversarial training. Specifically, we find that misclassified examples indeed have a significant impact on the final robustness. More surprisingly, we find that different maximization techniques on misclassified examples may have a negligible influence on the final robustness, while different minimization techniques are crucial. Motivated by the above discovery, we propose a new defense algorithm called Misclassification Aware adveRsarial Training (MART), which explicitly differentiates the misclassified and correctly classified examples during the training. We also propose a semi-supervised extension of MART, which can leverage the unlabeled data to further improve the robustness. Experimental results show that MART and its variant could significantly improve the state-of-the-art adversarial robustness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 66b3ff93-7dc4-459f-8ab0-8cb37811b742Cited by top-tier papers259
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
- Attacks Which Do Not Kill Training Make Adversarial Learning StrongerJingfeng Zhang, Xilie Xu, Bo Han, Gang Niu et al.ICML 2020 · 452 citations
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov et al.NeurIPS 2020 · 336 citations
Related papers
- Splitting the Difference on Adversarial TrainingMatan Levi, Aryeh KontorovichUSENIX Security 2024 · 9 citations
- Advancing Example Exploitation Can Alleviate Critical Challenges in Adversarial TrainingYao Ge, Yun Li, Keji Han, Junyi Zhu et al.ICCV 2023 · 6 citations
- Memorization Weights for Instance Reweighting in Adversarial TrainingJianfu Zhang, Yan Hong, Qibin ZhaoAAAI 2023 · 4 citations
- Are Adversarial Examples Created Equal? A Learnable Weighted Minimax Risk for Robustness under Non-uniform AttacksHuimin Zeng, Chen Zhu, Tom Goldstein, Furong HuangAAAI 2021 · 21 citations
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 38 citations
