Vision Transformers as Probabilistic Expansion from Learngene
Qiufeng Wang, Xu Yang, Haokun Chen, Xin Geng
摘要
We propose expanding the shared Transformer module to produce and initialize Transformers of varying depths, enabling adaptation to diverse resource constraints. Drawing an analogy to genetic expansibility, we term such module as learngene. To identify the expansion mechanism, we delve into the relationship between the layer's position and its corresponding weight value, and find that linear function appropriately approximates this relationship. Building on this insight, we present Transformer as Linear Expansion of learnGene (TLEG), a novel approach for flexibly producing and initializing Transformers of diverse depths. Specifically, to learn learngene, we firstly construct an auxiliary Transformer linearly expanded from learngene, after which we train it through employing soft distillation. Subsequently, we can produce and initialize Transformers of varying depths via linearly expanding the welltrained learngene, thereby supporting diverse downstream scenarios. Extensive experiments on ImageNet-1K demonstrate that TLEG achieves comparable or better performance in contrast to many individual models trained from scratch, while reducing around 2× training cost. When transferring to several downstream classification datasets, TLEG surpasses existing initialization methods by a large margin (e.g., +6.87% on iNat 2019 and +7.66% on CIFAR-100). Under the situation where we need to produce models of varying depths adapting for different resource constraints, TLEG achieves comparable results while reducing around 19× parameters stored to initialize these models and around 5× pre-training costs, in contrast to the pre-training and fine-tuning approach. When transferring a fixed set of parameters to initialize different models, TLEG presents better flexibility and competitive performance while reducing around 2.9× parameters stored to initialize, compared to the pre-training approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Cluster-Learngene: Inheriting Adaptive Clusters for Vision TransformersQiufeng Wang, Xu Yang, Fu Feng, Jing Wang 等NeurIPS 2024 · 被引用 10 次
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyQuanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen 等NeurIPS 2025 · 被引用 9 次
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 被引用 9 次
- Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale ModelsShuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang 等NeurIPS 2024 · 被引用 7 次
- Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitionsboyu shi, Chang Liu, Chuanbao Gao, Xu Yang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- Transformer as Linear Expansion of LearngeneShiyu Xia, Miaosen Zhang, Xu Yang, Ruiming Chen 等AAAI 2024 · 被引用 14 次
- Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic EnvironmentsShuxia Lin, Qiufeng Wang, Chang Liu, Xu Yang 等AAAI 2026
- WAVE: Weight Templates for Adaptive Initialization of Variable-sized ModelsFu Feng, Yucheng Xie, Jing Wang, Xin GengCVPR 2025
- Building Variable-Sized Models via Learngene PoolBoyu Shi, Shiyu Xia, Xu Yang, Haokun Chen 等AAAI 2024 · 被引用 5 次
- Inheriting Generalized Learngene for Efficient Knowledge Transfer across Multiple TasksYuankun Zu, Shiyu Xia, Xu Yang, Qiufeng Wang 等AAAI 2025
