Revisiting Token Dropping Strategy in Efficient BERT Pretraining
Qihuang Zhong, Liang Ding, Juhua Liu, Xuebo Liu, Min Zhang, Bo Du, Dacheng Tao
摘要
Token dropping is a recently-proposed strategy to speed up the pretraining of masked language models, such as BERT, by skipping the computation of a subset of the input tokens at several middle layers. It can effectively reduce the training time without degrading much performance on downstream tasks. However, we empirically find that token dropping is prone to a semantic loss problem and falls short in handling semantic-intense tasks ( §2). Motivated by this, we propose a simple yet effective semantic-consistent learning method (SCTD) to improve the token dropping. SCTD aims to encourage the model to learn how to preserve the semantic information in the representation space. Extensive experiments on 12 tasks show that, with the help of our SCTD, token dropping can achieve consistent and significant performance gains across all task types and model sizes. More encouragingly, SCTD saves up to 57% of pretraining time and brings up to +1.56% average improvement over the vanilla token dropping.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Not All Tokens Are What You Need for PretrainingZhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu 等NeurIPS 2024 · 被引用 99 次
- Accelerating Transformers with Spectrum-Preserving Token MergingChau Tran, Duy M. H. Nguyen, Manh-Duy Nguyen, TrungTin Nguyen 等NeurIPS 2024 · 被引用 51 次
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenCVPR 2026 · 被引用 18 次
- Nonparametric Teaching of Attention LearnersChen Zhang, Jianghui Wang, Bingyang Cheng, Zhongtao Chen 等ICLR 2026 · 被引用 3 次
- Unlocking Full Efficiency of Token Filtering in Large Language Model TrainingDi Chai, LI Pengbo, Feiyuan Zhang, Yilun Jin 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Self-Distillation as Instance-Specific Label SmoothingZhilu Zhang, Mert R. SabuncuNeurIPS 2020 · 被引用 155 次
相关 Paper
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu 等ACL 2022
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend 等ICLR 2021 · 被引用 83 次
- AdapLeR: Speeding up Inference by Adaptive Length ReductionAli Modarressi, Hosein Mohebbi, Mohammad Taher PilehvarACL 2022 · 被引用 34 次
- Scaling Language-Image Pre-Training via MaskingYanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer 等CVPR 2023
