LeeBERT: Learned Early Exit for BERT with cross-level optimization
Wei Zhu
Abstract
Pre-trained language models like BERT are performant in a wide range of natural language tasks. However, they are resource exhaustive and computationally expensive for industrial scenarios. Thus, early exits are adopted at each layer of BERT to perform adaptive computation by predicting easier samples with the first few layers to speed up the inference. In this work, to improve efficiency without performance drop, we propose a novel training scheme called Learned Early Exiting for BERT (LeeBERT). First, we ask each exit to learn from each other, rather than learning only from the last layer. Second, the weights of different loss terms are learned, thus balancing off different objectives. We formulate the optimization of LeeBERT as a bi-level optimization problem, and we propose a novel cross-level optimization (CLO) algorithm to improve the optimization results. Experiments on the GLUE benchmark show that our proposed methods improve the performance of the state-of-the-art (SOTA) early exiting methods for pre-trained models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani et al.NeurIPS 2022 · 394 citations
- Overthinking the Truth: Understanding how Language Models Process False DemonstrationsDanny Halawi, Jean-Stanislas Denain, Jacob SteinhardtICLR 2024 · 83 citations
- ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceZiqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang et al.AAAI 2024 · 26 citations
- PALBERT: Teaching ALBERT to PonderNikita Balagansky, Daniil GavrilovNeurIPS 2022 · 10 citations
- SPT: Learning to Selectively Insert Prompts for Better Prompt TuningWei Zhu, Ming TanEMNLP 2023 · 7 citations
Builds on10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
Related papers
- GAML-BERT: Improving BERT Early Exiting by Gradient Aligned Mutual LearningWei Zhu, Xiaoling Wang, Yuan Ni, Guotong XieEMNLP 2021 · 12 citations
- COST-EFF: Collaborative Optimization of Spatial and Temporal Efficiency with Slenderized Multi-exit Language ModelsBowen Shen, Zheng Lin, Yuanxin Liu, Zhengxiao Liu et al.EMNLP 2022 · 3 citations
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismJianing He, Qi Zhang, Hongyun Zhang, Xuanjing Huang et al.AAAI 2025 · 3 citations
- EarlyBERT: Efficient BERT Training via Early-bird Lottery TicketsXiaohan Chen, Yu Cheng, Shuohang Wang, Zhe Gan et al.ACL 2021
