SkipBERT: Efficient Inference with Shallow Layer Skipping
Jue Wang, Ke Chen, Gang Chen, Lidan Shou, Julian J. McAuley
Abstract
In this paper, we propose SkipBERT to accelerate BERT inference by skipping the computation of shallow layers. To achieve this, our approach encodes small text chunks into independent representations, which are then materialized to approximate the shallow representation of BERT. Since the use of such approximation is inexpensive compared with transformer calculations, we leverage it to replace the shallow layers of BERT to skip their runtime overhead. With off-the-shelf early exit mechanisms, we also skip redundant computation from the highest few layers to further improve inference efficiency. Results on GLUE show that our approach can reduce latency by 65% without sacrificing performance. By using only two-layer transformer calculations, we can still maintain 95% accuracy of BERT. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 569cee30-17a4-4d91-81bb-3d349aeb3627Cited by top-tier papers6
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 96 citations
- DeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyYixuan Weng, Minjun Zhu, Qiujie Xie, Qiyao Sun et al.ICLR 2026 · 57 citations
- ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceZiqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang et al.AAAI 2024 · 26 citations
- Retraining-free Model Quantization via One-Shot Weight-Coupling LearningChen Tang, Yuan Meng, Jiacheng Jiang, Shuzhao Xie et al.CVPR 2024 · 5 citations
- Routing Experts: Learning to Route Dynamic Experts in Existing Multi-modal Large Language ModelsQiong Wu, Zhaoxi Ke, Yiyi Zhou, Xiaoshuai Sun et al.ICLR 2025
Builds on8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley et al.NeurIPS 2020 · 473 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
Related papers
- Transkimmer: Transformer Learns to Layer-wise SkimYue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin et al.ACL 2022
- PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector EliminationSaurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy et al.ICML 2020 · 260 citations
- COST-EFF: Collaborative Optimization of Spatial and Temporal Efficiency with Slenderized Multi-exit Language ModelsBowen Shen, Zheng Lin, Yuanxin Liu, Zhengxiao Liu et al.EMNLP 2022 · 3 citations
- schuBERT: Optimizing Elements of BERTAshish Khetan, Zohar S. KarninACL 2020 · 2 citations
- Block-Skim: Efficient Question Answering for TransformerYue Guan, Zhengyi Li, Zhouhan Lin, Yuhao Zhu et al.AAAI 2022 · 33 citations
