Cluster-Learngene: Inheriting Adaptive Clusters for Vision Transformers
Qiufeng Wang, Xu Yang, Fu Feng, Jing Wang, Xin Geng
Abstract
In recent years, the merging of vast datasets with powerful computational resources has led to the emergence of large pre-trained models in the field of deep learning. However, the common practices often overgeneralize the applicability of these models, overlooking the task-specific resource constraints. To mitigate this issue, we propose Cluster-Learngene, which effectively clusters critical internal modules from a large ancestry model and then inherits them to initialize descendant models of elastic scales. Specifically, based on the density characteristics of attention heads, our method adaptively clusters attention heads of each layer and position-wise feed-forward networks (FFNs) in the ancestry model as the learngene. Moreover, we introduce priority weight-sharing and learnable parameter transformations that expand the learngene to initialize descendant models of elastic scales. Through extensive experimentation, we demonstrate that Cluster-Learngene not only is more efficient compared to other initialization methods but also customizes models of elastic scales according to downstream task resources.
† Corresponding authors. * The terms "foundation model" and "ancestry model," as well as "downstream model" and "descendant model," are interchangeably utilized unless distinctions are explicitly mentioned.
38th Conference on Neural Information Processing Systems (NeurIPS 2024).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 032ec65f-309b-4252-96cf-95d4be17cca6Cited by top-tier papers6
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene ConsistencyQuanjian Song, Donghao Zhou, Jingyu Lin, Fei Shen et al.NeurIPS 2025 · 9 citations
- ECO: Evolving Core Knowledge for Efficient TransferFu Feng, Yucheng Xie, Ruixiao Shi, Jianlu Shen et al.NeurIPS 2025 · 4 citations
- WAVE: Weight Templates for Adaptive Initialization of Variable-sized ModelsFu Feng, Yucheng Xie, Jing Wang, Xin GengCVPR 2025
- Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic EnvironmentsShuxia Lin, Qiufeng Wang, Chang Liu, Xu Yang et al.AAAI 2026
- LIVE: Learnable In-Context Vector for Visual Question AnsweringYingzhe Peng, Chenduo Hao, Xinting Hu, Jiawei Peng et al.NeurIPS 2024
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
Related papers
- Vision Transformers as Probabilistic Expansion from LearngeneQiufeng Wang, Xu Yang, Haokun Chen, Xin GengICML 2024 · 6 citations
- Transformer as Linear Expansion of LearngeneShiyu Xia, Miaosen Zhang, Xu Yang, Ruiming Chen et al.AAAI 2024 · 14 citations
- LEMON: Lossless model expansionYite Wang, Jiahao Su, Hanlin Lu, Cong Xie et al.ICLR 2024 · 25 citations
- Initializing Variable-sized Vision Transformers from Learngene with Learnable TransformationShiyu Xia, Yuankun Zu, Xu Yang, Xin GengNeurIPS 2024 · 9 citations
- FINE: Factorizing Knowledge for Initialization of Variable-sized Diffusion ModelsYucheng Xie, Fu Feng, Ruixiao Shi, Jianlu Shen et al.CVPR 2026
