Dynamic Transformers Provide a False Sense of Efficiency
Yiming Chen, Simin Chen, Zexin Li, Wei Yang, Cong Liu, Robby T. Tan, Haizhou Li
Abstract
Despite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference. Multi-exit is a mainstream approach to address this issue by making a trade-off between efficiency and accuracy, where the saving of computation comes from an early exit. However, whether such saving from early-exiting is robust remains unknown. Motivated by this, we first show that directly adapting existing adversarial attack approaches targeting model accuracy cannot significantly reduce inference efficiency. To this end, we propose a simple yet effective attacking framework, SAME, a novel slowdown attack framework on multi-exit models, which is specially tailored to reduce the efficiency of the multi-exit models. By leveraging the multi-exit models' design characteristics, we utilize all internal predictions to guide the adversarial sample generation instead of merely considering the final prediction. Experiments on the GLUE benchmark show that SAME can effectively diminish the efficiency gain of various multi-exit models by 80% on average, convincingly validating its effectiveness and generalization ability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc361f61-7f86-4270-a0b2-52e21b60b04eCited by top-tier papers7
- That Doesn't Go There: Attacks on Shared State in Multi-User Augmented Reality ApplicationsCarter Slocum, Yicheng Zhang, Erfan Shayegani, Pedram Zaree et al.USENIX Security 2024 · 21 citations
- LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless LoopsJiyuan Fu, Kaixun Jiang, Lingyi Hong, Jinglun Li et al.ICLR 2026 · 12 citations
- RT-LM: Uncertainty-Aware Resource Management for Real-Time Inference of Language ModelsYufei Li, Zexin Li, Wei Yang, Cong LiuRTSS 2023 · 10 citations
- : On-Device Real-Time Deep Reinforcement Learning for Autonomous RoboticsZexin Li, Aritra Samanta, Yufei Li, Andrea Soltoggio et al.RTSS 2023 · 9 citations
- RED: A Systematic Real-Time Scheduling Approach for Robotic Environmental DynamicsZexin Li, Tao Ren, Xiaoxi He, Cong LiuRTSS 2023 · 8 citations
Builds on18
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
Related papers
- BERT Lost Patience Won't Be Robust to Adversarial SlowdownZachary Coalson, Gabriel Ritter, Rakesh Bobba, Sanghyun HongNeurIPS 2023 · 7 citations
- COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismJianing He, Qi Zhang, Hongyun Zhang, Xuanjing Huang et al.AAAI 2025 · 3 citations
- A Panda? No, It's a Sloth: Slowdown Attacks on Adaptive Multi-Exit Neural Network InferenceSanghyun Hong, Yigitcan Kaya, Ionut-Vlad Modoranu, Tudor DumitrasICLR 2021 · 85 citations
- NICGSlowDown: Evaluating the Efficiency Robustness of Neural Image Caption Generation ModelsSimin Chen, Zihe Song, Mirazul Haque, Cong Liu et al.CVPR 2022 · 34 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
