Inequality phenomenon in l∞-adversarial training, and its unrealized threats
Ranjie Duan, Yuefeng Chen, Yao Zhu, Xiaojun Jia, Rong Zhang, Hui Xue
Abstract
The appearance of adversarial examples raises attention from both academia and industry. Along with the attack-defense arms race, adversarial training is the most effective against adversarial examples.However, we find inequality phenomena occur during the -adversarial training, that few features dominate the prediction made by the adversarially trained model. We systematically evaluate such inequality phenomena by extensive experiments and find such phenomena become more obvious when performing adversarial training with increasing adversarial strength (evaluated by ). We hypothesize such inequality phenomena make -adversarially trained model less reliable than the standard trained model when few ``important features" are influenced. To validate our hypothesis, we proposed two simple attacks that either perturb or replace important features with noise or occlusion. Experiments show that -adversarially trained model can be easily attacked when the few important features are influenced. Our work shed light on the limitation of the practicality of -adversarial training.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cb9f0bb4-67ab-49e6-a819-e6b8c0077ab9Related papers
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen AttacksDavid Stutz, Matthias Hein, Bernt SchieleICML 2020 · 158 citations
- More Data Can Expand The Generalization Gap Between Adversarially Robust and Standard ModelsLin Chen, Yifei Min, Mingrui Zhang, Amin KarbasiICML 2020 · 66 citations
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang et al.ICLR 2024
- Explicit Tradeoffs between Adversarial and Natural Distributional RobustnessMazda Moayeri, Kiarash Banihashem, Soheil FeiziNeurIPS 2022 · 28 citations
