bert2BERT: Towards Reusable Pretrained Language Models
Cheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu
摘要
In recent years, researchers tend to pre-train ever-larger language models to explore the upper limit of deep models. However, large language model pre-training costs intensive computational resources, and most of the models are trained from scratch without reusing the existing pre-trained models, which is wasteful. In this paper, we propose bert2BERT, which can effectively transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and significantly improve the pre-training efficiency of the large model. Specifically, we extend the previous function-preserving method proposed in computer vision on the Transformer-based language model, and further improve it by proposing a novel method, advanced knowledge for large model’s initialization. In addition, a two-stage learning method is proposed to further accelerate the pre-training. We conduct extensive experiments on representative PLMs (e.g., BERT and GPT) and demonstrate that (1) our method can save a significant amount of training cost compared with baselines including learning from scratch, StackBERT and MSLT; (2) our method is generic and applicable to different types of pre-trained models. In particular, bert2BERT saves about 45% and 47% computational cost of pre-training BERT_BASE and GPT_BASE by reusing the models of almost their half sizes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Masked Structural Growth for 2x Faster Language Model Pre-trainingYiqun Yao, Zheng Zhang, Jing Li, Yequan WangICLR 2024 · 被引用 30 次
- On the Inductive Bias of Stacking Towards Improving ReasoningNikunj Saunshi, Stefani Karp, Shankar Krishnan, Sobhan Miryoosefi 等NeurIPS 2024 · 被引用 23 次
- Reusing Pretrained Models by Multi-linear Operators for Efficient TrainingYu Pan, Ye Yuan, Yichun Yin, Zenglin Xu 等NeurIPS 2023 · 被引用 23 次
- Seeking Neural Nuggets: Knowledge Transfer in Large Language Models from a Parametric PerspectiveMing Zhong, Chenxin An, Weizhu Chen, Jiawei Han 等ICLR 2024 · 被引用 16 次
- Preparing Lessons for Progressive Training on Language ModelsYu Pan, Ye Yuan, Yichun Yin, Jiaxin Shi 等AAAI 2024 · 被引用 14 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
相关 Paper
- Accelerating Training of Transformer-Based Language Models with Progressive Layer DroppingMinjia Zhang, Yuxiong HeNeurIPS 2020 · 被引用 126 次
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao 等ACL 2025 · 被引用 6 次
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu 等ACL 2022
- LEMON: Lossless model expansionYite Wang, Jiahao Su, Hanlin Lu, Cong Xie 等ICLR 2024 · 被引用 25 次
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang 等ICLR 2022 · 被引用 23 次
