On the Vulnerability of Adversarially Trained Models Against Two-faced Attacks
Shengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang, Bo An, Lei Feng
摘要
Adversarial robustness is an important standard for measuring the quality of learned models, and adversarial training is an effective strategy for improving the adversarial robustness of models. In this paper, we disclose that adversarially trained models are vulnerable to two-faced attacks, where slight perturbations in input features are crafted to make the model exhibit a false sense of robustness in the verification phase. Such a threat is significantly important as it can mislead our evaluation of the adversarial robustness of models, which could cause unpredictable security issues when deploying substandard models in reality. More seriously, this threat seems to be pervasive and tricky: we find that many types of models suffer from this threat, and models with higher adversarial robustness tend to be more vulnerable. Furthermore, we provide the first attempt to formulate this threat, disclose its relationships with adversarial risk, and try to circumvent it via a simple countermeasure. These findings serve as a crucial reminder for practitioners to exercise caution in the verification phase, urging them to refrain from blindly trusting the exhibited adversarial robustness of models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
相关 Paper
- With False Friends Like These, Who Can Notice Mistakes?Lue Tao, Lei Feng, Jinfeng Yi, Songcan ChenAAAI 2022 · 被引用 6 次
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang 等NeurIPS 2021 · 被引用 90 次
- Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured DataBinghui Li, Yuanzhi LiICLR 2025
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot 等ICML 2020 · 被引用 103 次
- Adversarial Feature DesensitizationPouya Bashivan, Reza Bayat, Adam Ibrahim, Kartik Ahuja 等NeurIPS 2021 · 被引用 22 次
