Revisiting Token Dropping Strategy in Efficient BERT Pretraining
Qihuang Zhong, Liang Ding, Juhua Liu, Xuebo Liu, Min Zhang, Bo Du, Dacheng Tao
Abstract
Token dropping is a recently-proposed strategy to speed up the pretraining of masked language models, such as BERT, by skipping the computation of a subset of the input tokens at several middle layers. It can effectively reduce the training time without degrading much performance on downstream tasks. However, we empirically find that token dropping is prone to a semantic loss problem and falls short in handling semantic-intense tasks ( §2). Motivated by this, we propose a simple yet effective semantic-consistent learning method (SCTD) to improve the token dropping. SCTD aims to encourage the model to learn how to preserve the semantic information in the representation space. Extensive experiments on 12 tasks show that, with the help of our SCTD, token dropping can achieve consistent and significant performance gains across all task types and model sizes. More encouragingly, SCTD saves up to 57% of pretraining time and brings up to +1.56% average improvement over the vanilla token dropping.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9cbb277-110d-4701-bb79-ec6bc1fd9ae7Cited by top-tier papers9
- Not All Tokens Are What You Need for PretrainingZhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu et al.NeurIPS 2024 · 99 citations
- Accelerating Transformers with Spectrum-Preserving Token MergingChau Tran, Duy M. H. Nguyen, Manh-Duy Nguyen, TrungTin Nguyen et al.NeurIPS 2024 · 51 citations
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenCVPR 2026 · 18 citations
- Nonparametric Teaching of Attention LearnersChen Zhang, Jianghui Wang, Bingyang Cheng, Zhongtao Chen et al.ICLR 2026 · 3 citations
- Unlocking Full Efficiency of Token Filtering in Large Language Model TrainingDi Chai, LI Pengbo, Feiyuan Zhang, Yilun Jin et al.ICLR 2026 · 2 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Self-Distillation as Instance-Specific Label SmoothingZhilu Zhang, Mert R. SabuncuNeurIPS 2020 · 155 citations
Related papers
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 126 citations
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend et al.ICLR 2021 · 83 citations
- AdapLeR: Speeding up Inference by Adaptive Length ReductionAli Modarressi, Hosein Mohebbi, Mohammad Taher PilehvarACL 2022 · 34 citations
- Scaling Language-Image Pre-Training via MaskingYanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer et al.CVPR 2023
