BERT Lost Patience Won't Be Robust to Adversarial Slowdown
Zachary Coalson, Gabriel Ritter, Rakesh Bobba, Sanghyun Hong
Abstract
In this paper, we systematically evaluate the robustness of multi-exit language models against adversarial slowdown. To audit their robustness, we design a slowdown attack that generates natural adversarial text bypassing early-exit points. We use the resulting WAFFLE attack as a vehicle to conduct a comprehensive evaluation of three multi-exit mechanisms with the GLUE benchmark against adversarial slowdown. We then show our attack significantly reduces the computational savings provided by the three methods in both white-box and black-box settings. The more complex a mechanism is, the more vulnerable it is to adversarial slowdown. We also perform a linguistic analysis of the perturbed text inputs, identifying common perturbation patterns that our attack generates, and comparing them with standard adversarial text attacks. Moreover, we show that adversarial training is ineffective in defeating our slowdown attack, but input sanitization with a conversational model, e.g., ChatGPT, can remove perturbations effectively. This result suggests that future work is needed for developing efficient yet robust multi-exit models. Our code is available at: https://github.com/ztcoalson/WAFFLE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c93c8a80-d8fb-410d-82fb-f0a3b8ea0527Cited by top-tier papers2
- SoK: Efficiency Robustness of Dynamic Deep Learning SystemsRavishka Rathnasuriya, Tingxi Li, Zexin Xu, Zihe Song et al.USENIX Security 2025
- NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language ModelsYiming Chen, Zexin Li, Xianghu Yue, Robby T. Tan et al.ACL 2026
Builds on15
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
Related papers
- Dynamic Transformers Provide a False Sense of EfficiencyYiming Chen, Simin Chen, Zexin Li, Wei Yang et al.ACL 2023 · 5 citations
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui et al.ICLR 2024 · 146 citations
- Generative Adversarial Training with Perturbed Token Detection for Model RobustnessJiahao Zhao, Wenji MaoEMNLP 2023 · 3 citations
- SAFENUDGE: Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offsJoão Fonseca, Andrew Bell, Julia StoyanovichEMNLP 2025 · 1 citation
- White-Box Multi-Objective Adversarial Attack on Dialogue GenerationYufei Li, Zexin Li, Yingfan Gao, Cong LiuACL 2023 · 12 citations
