WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models
Fu Feng, Yucheng Xie, Jing Wang, Xin Geng
Abstract
The growing complexity of model parameters underscores the significance of pre-trained models. However, deployment constraints often necessitate models of varying sizes, exposing limitations in the conventional pre-training and fine-tuning paradigm, particularly when target model sizes are incompatible with pre-trained ones. To address this challenge, we propose WAVE, a novel approach that reformulates variable-sized model initialization from a multitask perspective, where initializing each model size is treated as a distinct task. WAVE employs shared, sizeagnostic weight templates alongside size-specific weight scalers to achieve consistent initialization across various model sizes. These weight templates, constructed within the Learngene framework, integrate knowledge from pretrained models through a distillation process constrained by Kronecker-based rules. Target models are then initialized by concatenating and weighting these templates, with adaptive connection rules established by lightweight weight scalers, whose parameters are learned from minimal training data. Extensive experiments demonstrate the efficiency of WAVE, achieving state-of-the-art performance in initializing models of various depth and width. The knowledge encapsulated in weight templates is also task-agnostic, allowing for seamless transfer across diverse downstream datasets. Code will be made available at https://github.com/fu-feng/WAVE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Stratified Knowledge-Density Super-Network for Scalable Vision TransformersLonghua Li, Lei Qi, Xin GengAAAI 2026 · 1 citation
- A Unified Framework for Knowledge Transfer in Bidirectional Model ScalingJianlu Shen, Fu Feng, Jiaze Xu, Yucheng Xie et al.CVPR 2026 · 1 citation
- FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion ModelsYucheng Xie, Fu Feng, Ruixiao Shi, Jianlu Shen et al.CVPR 2026
- HAP: Harmonized Amplitude Perturbation for Cross-Domain Few-Shot LearningWenqian Li, Pengfei Fang, Hui XueAAAI 2026
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
Related papers
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 9 citations
- Self-Supervised Weight Templates for Scalable Vision Model InitializationYucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang et al.ICML 2026 · 1 citation
- Vision Transformers as Probabilistic Expansion from LearngeneQiufeng Wang, Xu Yang, Haokun Chen, Xin GengICML 2024 · 6 citations
- Transformer as Linear Expansion of LearngeneShiyu Xia, Miaosen Zhang, Xu Yang, Ruiming Chen et al.AAAI 2024 · 14 citations
- Learngene Tells You How to Customize: Task-Aware Parameter Initialization at Flexible ScalesJiaze Xu, Shiyu Xia, Xu Yang, Jiaqi Lv et al.ICML 2025
