Can Adversarial Training Be Manipulated By Non-Robust Features?
Lue Tao, Lei Feng, Hongxin Wei, Jinfeng Yi, Sheng-Jun Huang, Songcan Chen
Abstract
Adversarial training, originally designed to resist test-time adversarial examples, has shown to be promising in mitigating training-time availability attacks. This defense ability, however, is challenged in this paper. We identify a novel threat model named stability attack, which aims to hinder robust availability by slightly manipulating the training data. Under this threat, we show that adversarial training using a conventional defense budget provably fails to provide test robustness in a simple statistical setting, where the non-robust features of the training data can be reinforced by -bounded perturbation. Further, we analyze the necessity of enlarging the defense budget to counter stability attacks. Finally, comprehensive experiments demonstrate that stability attacks are harmful on benchmark datasets, and thus the adaptive defense is necessary to maintain robustness. Our code is available at https://github.com/TLMichael/Hypocritical-Perturbation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e658655-a113-439a-a647-dde34045a62aCited by top-tier papers9
- Adversarial Examples Are Not Real FeaturesAng Li, Yifei Wang, Yiwen Guo, Yisen WangNeurIPS 2023 · 24 citations
- Efficient Availability Attacks against Supervised and Contrastive Learning SimultaneouslyYihan Wang, Yifan Zhu, Xiao-Shan GaoNeurIPS 2024 · 14 citations
- Toward Availability Attacks in 3D Point CloudsYifan Zhu, Yibo Miao, Yinpeng Dong, Xiao-Shan GaoICML 2024 · 9 citations
- Theoretical Understanding of Learning from Adversarial PerturbationsSoichiro Kumano, Hiroshi Kera, Toshihiko YamasakiICLR 2024 · 4 citations
- What Distributions are Robust to Indiscriminate Poisoning Attacks for Linear Learners?Fnu Suya, Xiao Zhang, Yuan Tian, David EvansNeurIPS 2023 · 3 citations
Builds on32
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot et al.ICML 2020 · 103 citations
- On Robustness of Linear Classifiers to Targeted Data PoisoningNakshatra Gupta, Sumanth Prabhu S, Supratik Chakraborty, R. VenkateshAAAI 2026
- Why adversarial training can hurt robust accuracyJacob Clarysse, Julia Hörrmann, Fanny YangICLR 2023 · 6 citations
- Efficient Adversarial Training With Transferable Adversarial ExamplesHaizhong Zheng, Ziqi Zhang, Juncheng Gu, Honglak Lee et al.CVPR 2020
