SkipBERT: Efficient Inference with Shallow Layer Skipping
Jue Wang, Ke Chen, Gang Chen, Lidan Shou, Julian J. McAuley
摘要
In this paper, we propose SkipBERT to accelerate BERT inference by skipping the computation of shallow layers. To achieve this, our approach encodes small text chunks into independent representations, which are then materialized to approximate the shallow representation of BERT. Since the use of such approximation is inexpensive compared with transformer calculations, we leverage it to replace the shallow layers of BERT to skip their runtime overhead. With off-the-shelf early exit mechanisms, we also skip redundant computation from the highest few layers to further improve inference efficiency. Results on GLUE show that our approach can reduce latency by 65% without sacrificing performance. By using only two-layer transformer calculations, we can still maintain 95% accuracy of BERT. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A Survey on Model Compression and Acceleration for Pretrained Language ModelsCanwen Xu, Julian J. McAuleyAAAI 2023 · 被引用 96 次
- DeepScientist: Advancing Frontier-Pushing Scientific Findings ProgressivelyYixuan Weng, Minjun Zhu, Qiujie Xie, Qiyao Sun 等ICLR 2026 · 被引用 57 次
- ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models InferenceZiqian Zeng, Yihuai Hong, Hongliang Dai, Huiping Zhuang 等AAAI 2024 · 被引用 26 次
- Retraining-free Model Quantization via One-Shot Weight-Coupling LearningChen Tang, Yuan Meng, Jiacheng Jiang, Shuzhao Xie 等CVPR 2024 · 被引用 5 次
- Routing Experts: Learning to Route Dynamic Experts in Existing Multi-modal Large Language ModelsQiong Wu, Zhaoxi Ke, Yiyi Zhou, Xiaoshuai Sun 等ICLR 2025
它引用的顶会 Paper8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- BERT Loses Patience: Fast and Robust Inference with Early ExitWangchunshu Zhou, Canwen Xu, Tao Ge, Julian J. McAuley 等NeurIPS 2020 · 被引用 473 次
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao 等ACL 2020 · 被引用 257 次
相关 Paper
- Transkimmer: Transformer Learns to Layer-wise SkimYue Guan, Zhengyi Li, Jingwen Leng, Zhouhan Lin 等ACL 2022
- PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector EliminationSaurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy 等ICML 2020 · 被引用 260 次
- COST-EFF: Collaborative Optimization of Spatial and Temporal Efficiency with Slenderized Multi-exit Language ModelsBowen Shen, Zheng Lin, Yuanxin Liu, Zhengxiao Liu 等EMNLP 2022 · 被引用 3 次
- schuBERT: Optimizing Elements of BERTAshish Khetan, Zohar S. KarninACL 2020 · 被引用 2 次
- Block-Skim: Efficient Question Answering for TransformerYue Guan, Zhengyi Li, Zhouhan Lin, Yuhao Zhu 等AAAI 2022 · 被引用 33 次
