GAML-BERT: Improving BERT Early Exiting by Gradient Aligned Mutual Learning
Wei Zhu, Xiaoling Wang, Yuan Ni, Guotong Xie
Abstract
In this work, we propose a novel framework, Gradient Aligned Mutual Learning BERT (GAML-BERT), for improving the early exiting of BERT. GAML-BERT's contributions are two-fold. We conduct a set of pilot experiments, which shows that mutual knowledge distillation between a shallow exit and a deep exit leads to better performances for both. From this observation, we use mutual learning to improve BERT's early exiting performances, that is, we ask each exit of a multi-exit BERT to distill knowledge from each other. Second, we propose GA, a novel training method that aligns the gradients from knowledge distillation to cross-entropy losses. Extensive experiments are conducted on the GLUE benchmark, which shows that our GAML-BERT can significantly outperform the state-of-the-art (SOTA) BERT early exiting methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6f31f29-1226-4b1a-8928-bf509e4247a9Cited by top-tier papers4
- SPT: Learning to Selectively Insert Prompts for Better Prompt TuningWei Zhu, Ming TanEMNLP 2023 · 7 citations
- COSEE: Consistency-Oriented Signal-Based Early Exiting via Calibrated Sample Weighting MechanismJianing He, Qi Zhang, Hongyun Zhang, Xuanjing Huang et al.AAAI 2025 · 3 citations
- RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient InferenceLianming Huang, Shangyu Wu, Yufei Cui, Ying Xiong et al.ICLR 2026 · 3 citations
- EnViT: Enhancing the Performance of Early-Exit Vision Transformers via Exit-Aware Structured Dropout-Enabled Self-DistillationYonghao Dong, Qiang He, Penghong Rui, Zhenzhe Zheng et al.AAAI 2026
Builds on9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
- Distillation-Based Training for Multi-Exit ArchitecturesMary Phuong, Christoph LampertICCV 2019 · 205 citations
Related papers
- LeeBERT: Learned Early Exit for BERT with cross-level optimizationWei ZhuACL 2021
- COST-EFF: Collaborative Optimization of Spatial and Temporal Efficiency with Slenderized Multi-exit Language ModelsBowen Shen, Zheng Lin, Yuanxin Liu, Zhengxiao Liu et al.EMNLP 2022 · 3 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
- SKDBERT: Compressing BERT via Stochastic Knowledge DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo et al.AAAI 2023 · 13 citations
- How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo et al.AAAI 2024 · 4 citations
