Initializing Variable-sized Vision Transformers from Learngene with Learnable Transformation
Shiyu Xia, Yuankun Zu, Xu Yang, Xin Geng
摘要
In practical scenarios, it is necessary to build variable-sized models to accommodate diverse resource constraints, where weight initialization serves as a crucial step preceding training. The recently introduced Learngene framework firstly learns one compact module, termed learngene , from a large well-trained model, and then transforms learngene to initialize variable-sized models. However, the existing Learngene methods provide limited guidance on transforming learngene, where transformation mechanisms are manually designed and generally lack a learnable component. Moreover, these methods only consider transforming learngene along depth dimension, thus constraining the flexibility of learngene. Motivated by these concerns, we propose a novel and effective Learngene approach termed LeTs ( Le arnable T ran s formation ), where we transform the learngene module along both width and depth dimension with a set of learnable matrices for flexible variable-sized model initialization. Specifically, we construct an auxiliary model comprising the compact learngene module and learnable transformation matrices, enabling both components to be trained. To meet the varying size requirements of target models, we select specific parameters from well-trained transformation matrices to adaptively transform the learngene, guided by strategies such as continuous selection and magnitude-wise selection. Extensive experiments on ImageNet-1K demonstrate that Des-Nets initialized via LeTs outperform those with 100-epoch from scratch training after only 1 epoch tuning. When transferring to downstream image classification tasks, LeTs achieves better results while outperforming from scratch training after about 10 epochs within a 300-epoch training schedule.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- WAVE: Weight Templates for Adaptive Initialization of Variable-sized ModelsFu Feng, Yucheng Xie, Jing Wang, Xin GengCVPR 2025
- When Labelers Stay Silent: The Power of Ties in Cost-Effective Preference LearningJIAQI LYU, Zihan Zhang, CJ Y, Shiyu Xia 等ICML 2026
- Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation ModelsJiawei Fan, Shigeng Wang, Chao Li, Xiaolong Liu 等CVPR 2026
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
它引用的顶会 Paper37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceBenjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock 等ICCV 2021 · 被引用 1,009 次
相关 Paper
- Transformer as Linear Expansion of LearngeneShiyu Xia, Miaosen Zhang, Xu Yang, Ruiming Chen 等AAAI 2024 · 被引用 14 次
- Vision Transformers as Probabilistic Expansion from LearngeneQiufeng Wang, Xu Yang, Haokun Chen, Xin GengICML 2024 · 被引用 6 次
- Building Variable-Sized Models via Learngene PoolBoyu Shi, Shiyu Xia, Xu Yang, Haokun Chen 等AAAI 2024 · 被引用 5 次
- Inheriting Generalized Learngene for Efficient Knowledge Transfer across Multiple TasksYuankun Zu, Shiyu Xia, Xu Yang, Qiufeng Wang 等AAAI 2025
- Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic EnvironmentsShuxia Lin, Qiufeng Wang, Chang Liu, Xu Yang 等AAAI 2026
