Reusing Pretrained Models by Multi-linear Operators for Efficient Training
Yu Pan, Ye Yuan, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang, Qun Liu
Abstract
Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12.0% and LiGO by +20.7%, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdb48872-eb44-4fd4-b6d1-c854fc6ff674Cited by top-tier papers9
- Preparing Lessons for Progressive Training on Language ModelsYu Pan, Ye Yuan, Yichun Yin, Jiaxin Shi et al.AAAI 2024 · 14 citations
- Measuring Vision-Language STEM Skills of Neural ModelsJianhao Shen, Ye Yuan, Srbuhi Mirzoyan, Ming Zhang et al.ICLR 2024 · 14 citations
- Benchmarking Ultra-Low-Power μNPUsJosh Millar, Yushan Huang, Sarab S. Sethi, Hamed Haddadi et al.MobiCom 2025 · 13 citations
- Deep Graph MatingYongcheng Jing, Seok-Hee Hong, Dacheng TaoNeurIPS 2024 · 9 citations
- SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive LearningQifan Yu, Xinyu Ma, Zhijian Zhuo, Minrui Wang et al.ICML 2026 · 3 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- bert2BERT: Towards Reusable Pretrained Language ModelsCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang et al.ACL 2022
- A Unified Framework for Knowledge Transfer in Bidirectional Model ScalingJianlu Shen, Fu Feng, Jiaze Xu, Yucheng Xie et al.CVPR 2026 · 1 citation
- A Multi-Level Framework for Accelerating Training Transformer ModelsLongwei Zou, Han Zhang, Yangdong DengICLR 2024 · 3 citations
- Late-to-Early Training: LET LLMs Learn Earlier, So Faster and BetterJi Zhao, Shitong Shao, Yufei Gu, Xun Zhou et al.ICLR 2026 · 1 citation
- LESA: Learnable LLM Layer Scaling-UpYifei Yang, Zouying Cao, Xinbei Ma, Yao Yao et al.ACL 2025 · 6 citations
