Extreme Miscalibration and the Illusion of Adversarial Robustness
Vyas Raina, Samson Tan, Volkan Cevher, Aditya Rawal, Sheng Zha, George Karypis
摘要
Deep learning-based Natural Language Processing (NLP) models are vulnerable to adversarial attacks, where small perturbations can cause a model to misclassify. Adversarial Training (AT) is often used to increase model robustness. However, we have discovered an intriguing phenomenon: deliberately or accidentally miscalibrating models masks gradients in a way that interferes with adversarial attack search methods, giving rise to an apparent increase in robustness. We show that this observed gain in robustness is an illusion of robustness (IOR), and demonstrate how an adversary can perform various forms of test-time temperature calibration to nullify the aforementioned interference and allow the adversarial attack to find adversarial examples. Hence, we urge the NLP community to incorporate test-time temperature scaling into their robustness evaluations to ensure that any observed gains are genuine. Finally, we show how the temperature can be scaled during training to improve genuine robustness. * Equal contribution. This paper reports on the work VR did whilst at Amazon Web Services. VC holds concurrent appointments as an Amazon Scholar and as a faculty at EPFL. This paper describes the work performed at Amazon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MixAT: Combining Continuous and Discrete Adversarial Training for LLMsCsaba Dékány, Stefan Balauca, Dimitar I. Dimitrov, Robin Staab 等NeurIPS 2025 · 被引用 12 次
- Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold PurificationChenhao Dang, Jing MaAAAI 2026
- Confidence Elicitation: A New Attack Vector for Large Language ModelsBrian Formento, Chuan-Sheng Foo, See-Kiong NgICLR 2025
它引用的顶会 Paper13
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun 等ICLR 2020 · 被引用 502 次
相关 Paper
- Enhance the Visual Representation via Discrete Adversarial TrainingXiaofeng Mao, Yuefeng Chen, Ranjie Duan, Yao Zhu 等NeurIPS 2022 · 被引用 48 次
- Improved Gradient-Based Adversarial Attacks for Quantized NetworksKartik Gupta, Thalaiyasingam AjanthanAAAI 2022 · 被引用 22 次
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang 等ICLR 2024
- How do humans perceive adversarial text? A reality check on the validity and naturalness of word-based adversarial attacksSalijona Dyrmishi, Salah Ghamizi, Maxime CordyACL 2023 · 被引用 4 次
- Analysis and Applications of Class-wise Robustness in Adversarial TrainingQi Tian, Kun Kuang, Kelu Jiang, Fei Wu 等KDD 2021 · 被引用 30 次
