With False Friends Like These, Who Can Notice Mistakes?
Lue Tao, Lei Feng, Jinfeng Yi, Songcan Chen
摘要
Adversarial examples crafted by an explicit adversary have attracted significant attention in machine learning. However, the security risk posed by a potential false friend has been largely overlooked. In this paper, we unveil the threat of hypocritical examples---inputs that are originally misclassified yet perturbed by a false friend to force correct predictions. While such perturbed examples seem harmless, we point out for the first time that they could be maliciously used to conceal the mistakes of a substandard (i.e., not as good as required) model during an evaluation. Once a deployer trusts the hypocritical performance and applies the "well-performed" model in real-world applications, unexpected failures may happen even in benign environments. More seriously, this security risk seems to be pervasive: we find that many types of substandard models are vulnerable to hypocritical examples across multiple datasets. Furthermore, we provide the first attempt to characterize the threat with a metric called hypocritical risk and try to circumvent it via several countermeasures. Results demonstrate the effectiveness of the countermeasures, while the risk remains non-negligible even after adaptive robust training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang 等NeurIPS 2021 · 被引用 90 次
- Can Adversarial Training Be Manipulated By Non-Robust Features?Lue Tao, Lei Feng, Hongxin Wei, Jinfeng Yi 等NeurIPS 2022 · 被引用 20 次
- Squeeze Training for Adversarial RobustnessQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenICLR 2023 · 被引用 2 次
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang 等ICLR 2024
它引用的顶会 Paper19
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
相关 Paper
- Doppelgangers and Adversarial VulnerabilityGeorge I. KamberovCVPR 2025
- Two Coupled Rejection Metrics Can Tell Adversarial Examples ApartTianyu Pang, Huishuai Zhang, Di He, Yinpeng Dong 等CVPR 2022 · 被引用 13 次
- Confidential Guardian: Cryptographically Prohibiting the Abuse of Model AbstentionStephan Rabanser, Ali Shahin Shamsabadi, Olive Franzese, Xiao Wang 等ICML 2025
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot 等ICML 2020 · 被引用 103 次
- Group-based Robustness: A General Framework for Customized Robustness in the Real WorldWeiran Lin, Keane Lucas, Neo Eyal, Lujo Bauer 等NDSS 2024
