WAVE: Weight Templates for Adaptive Initialization of Variable-sized Models
Fu Feng, Yucheng Xie, Jing Wang, Xin Geng
摘要
The growing complexity of model parameters underscores the significance of pre-trained models. However, deployment constraints often necessitate models of varying sizes, exposing limitations in the conventional pre-training and fine-tuning paradigm, particularly when target model sizes are incompatible with pre-trained ones. To address this challenge, we propose WAVE, a novel approach that reformulates variable-sized model initialization from a multitask perspective, where initializing each model size is treated as a distinct task. WAVE employs shared, sizeagnostic weight templates alongside size-specific weight scalers to achieve consistent initialization across various model sizes. These weight templates, constructed within the Learngene framework, integrate knowledge from pretrained models through a distillation process constrained by Kronecker-based rules. Target models are then initialized by concatenating and weighting these templates, with adaptive connection rules established by lightweight weight scalers, whose parameters are learned from minimal training data. Extensive experiments demonstrate the efficiency of WAVE, achieving state-of-the-art performance in initializing models of various depth and width. The knowledge encapsulated in weight templates is also task-agnostic, allowing for seamless transfer across diverse downstream datasets. Code will be made available at https://github.com/fu-feng/WAVE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Stratified Knowledge-Density Super-Network for Scalable Vision TransformersLonghua Li, Lei Qi, Xin GengAAAI 2026 · 被引用 1 次
- A Unified Framework for Knowledge Transfer in Bidirectional Model ScalingJianlu Shen, Fu Feng, Jiaze Xu, Yucheng Xie 等CVPR 2026 · 被引用 1 次
- FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion ModelsYucheng Xie, Fu Feng, Ruixiao Shi, Jianlu Shen 等CVPR 2026
- HAP: Harmonized Amplitude Perturbation for Cross-Domain Few-Shot LearningWenqian Li, Pengfei Fang, Hui XueAAAI 2026
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
相关 Paper
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 被引用 9 次
- Self-Supervised Weight Templates for Scalable Vision Model InitializationYucheng Xie, Fu Feng, Ruixiao Shi, Jing Wang 等ICML 2026 · 被引用 1 次
- Vision Transformers as Probabilistic Expansion from LearngeneQiufeng Wang, Xu Yang, Haokun Chen, Xin GengICML 2024 · 被引用 6 次
- Transformer as Linear Expansion of LearngeneShiyu Xia, Miaosen Zhang, Xu Yang, Ruiming Chen 等AAAI 2024 · 被引用 14 次
- Learngene Tells You How to Customize: Task-Aware Parameter Initialization at Flexible ScalesJiaze Xu, Shiyu Xia, Xu Yang, Jiaqi Lv 等ICML 2025
