Extreme Miscalibration and the Illusion of Adversarial Robustness
Vyas Raina, Samson Tan, Volkan Cevher, Aditya Rawal, Sheng Zha, George Karypis
Abstract
Deep learning-based Natural Language Processing (NLP) models are vulnerable to adversarial attacks, where small perturbations can cause a model to misclassify. Adversarial Training (AT) is often used to increase model robustness. However, we have discovered an intriguing phenomenon: deliberately or accidentally miscalibrating models masks gradients in a way that interferes with adversarial attack search methods, giving rise to an apparent increase in robustness. We show that this observed gain in robustness is an illusion of robustness (IOR), and demonstrate how an adversary can perform various forms of test-time temperature calibration to nullify the aforementioned interference and allow the adversarial attack to find adversarial examples. Hence, we urge the NLP community to incorporate test-time temperature scaling into their robustness evaluations to ensure that any observed gains are genuine. Finally, we show how the temperature can be scaled during training to improve genuine robustness. * Equal contribution. This paper reports on the work VR did whilst at Amazon Web Services. VC holds concurrent appointments as an Amazon Scholar and as a faculty at EPFL. This paper describes the work performed at Amazon.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- MixAT: Combining Continuous and Discrete Adversarial Training for LLMsCsaba Dékány, Stefan Balauca, Dimitar I. Dimitrov, Robin Staab et al.NeurIPS 2025 · 12 citations
- Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold PurificationChenhao Dang, Jing MaAAAI 2026
- Confidence Elicitation: A New Attack Vector for Large Language ModelsBrian Formento, Chuan-Sheng Foo, See-Kiong NgICLR 2025
Builds on13
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
Related papers
- Enhance the Visual Representation via Discrete Adversarial TrainingXiaofeng Mao, Yuefeng Chen, Ranjie Duan, Yao Zhu et al.NeurIPS 2022 · 48 citations
- Improved Gradient-Based Adversarial Attacks for Quantized NetworksKartik Gupta, Thalaiyasingam AjanthanAAAI 2022 · 22 citations
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang et al.ICLR 2024
- How do humans perceive adversarial text? A reality check on the validity and naturalness of word-based adversarial attacksSalijona Dyrmishi, Salah Ghamizi, Maxime CordyACL 2023 · 4 citations
- Analysis and Applications of Class-wise Robustness in Adversarial TrainingQi Tian, Kun Kuang, Kelu Jiang, Fei Wu et al.KDD 2021 · 30 citations
