LEMON: Lossless model expansion
Yite Wang, Jiahao Su, Hanlin Lu, Cong Xie, Tianyi Liu, Jianbo Yuan, Haibin Lin, Ruoyu Sun, Hongxia Yang
摘要
Scaling of deep neural networks, especially Transformers, is pivotal for their surging performance and has further led to the emergence of sophisticated reasoning capabilities in foundation models. Such scaling generally requires training large models from scratch with random initialization, failing to leverage the knowledge acquired by their smaller counterparts, which are already resource-intensive to obtain. To tackle this inefficiency, we present osslss del Expansio (LEMON), a recipe to initialize scaled models using the weights of their smaller but pre-trained counterparts. This is followed by model training with an optimized learning rate scheduler tailored explicitly for the scaled models, substantially reducing the training time compared to training from scratch. Notably, LEMON is versatile, ensuring compatibility with various network structures, including models like Vision Transformers and BERT. Our empirical results demonstrate that LEMON reduces computational costs by 56.7% for Vision Transformers and 33.2% for BERT when compared to training from scratch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SwitchTab: Switched Autoencoders Are Effective Tabular LearnersJing Wu, Suiyao Chen, Qi Zhao, Renat Sergazinov 等AAAI 2024 · 被引用 60 次
- On the Inductive Bias of Stacking Towards Improving ReasoningNikunj Saunshi, Stefani Karp, Shankar Krishnan, Sobhan Miryoosefi 等NeurIPS 2024 · 被引用 23 次
- Scaling depth capacity via zero/one-layer model expansionZhiqi BuICML 2026 · 被引用 1 次
- Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic TeacherYong Guo, Shulian Zhang, Haolin Pan, Jing Liu 等ICLR 2025
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
相关 Paper
- bert2BERT: Towards Reusable Pretrained Language ModelsCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang 等ACL 2022
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao 等ACL 2025 · 被引用 6 次
- LEMON: Reviving Stronger and Smaller LMs from Larger LMs with Linear Parameter FusionYilong Chen, Junyuan Shang, Zhenyu Zhang, Shiyao Cui 等ACL 2024 · 被引用 1 次
- FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model DeploymentRiccardo Zaccone, Stefanos Laskaridis, Marco Ciccone, Samuel HorváthICML 2026 · 被引用 1 次
- Learning to Grow Pretrained Models for Efficient Transformer TrainingPeihao Wang, Rameswar Panda, Lucas Torroba Hennigen, Philip Greengard 等ICLR 2023 · 被引用 13 次
