Linearly Decomposing and Recomposing Vision Transformers for Diverse-Scale Models
Shuxia Lin, Miaosen Zhang, Ruiming Chen, Xu Yang, Qiufeng Wang, Xin Geng
Abstract
Vision Transformers (ViTs) are widely used in a variety of applications, while they usually have a fixed architecture that may not match the varying computational resources of different deployment environments. Thus, it is necessary to adapt ViT architectures to devices with diverse computational overheads to achieve an accuracy-efficient trade-off. This concept is consistent with the motivation behind Learngene . To achieve this, inspired by polynomial decomposition in calculus, where a function can be approximated by linearly combining several basic components, we propose to linearly decompose the ViT model into a set of components called learngenes during element-wise training. These learngenes can then be recomposed into differently scaled, pre-initialized models to satisfy different computational resource constraints. Such a decomposition-recomposition strategy provides an economical and flexible approach to generating different scales of ViT models for different deployment scenarios. Compared to model compression or training from scratch, which require to repeatedly train on large datasets for diverse-scale models, such strategy reduces computational costs since it only requires to train on large datasets once. Extensive experiments are used to validate the effectiveness of our method: ViTs can be decomposed and the decomposed learngenes can be recomposed into diverse-scale ViTs, which can achieve comparable or better performance compared to traditional model compression and pre-training methods. The code for our experiments is available in the supplemental material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f37366db-ff4e-43ba-b3e6-247393c73e79Cited by top-tier papers8
- Cluster-Learngene: Inheriting Adaptive Clusters for Vision TransformersQiufeng Wang, Xu Yang, Fu Feng, Jing Wang et al.NeurIPS 2024 · 10 citations
- ECO: Evolving Core Knowledge for Efficient TransferFu Feng, Yucheng Xie, Ruixiao Shi, Jianlu Shen et al.NeurIPS 2025 · 4 citations
- FlexLoRA: Entropy-Guided Flexible Low-Rank AdaptationMuqing Liu, Chongjie Si, Yuheng JiaICLR 2026 · 3 citations
- Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic EnvironmentsShuxia Lin, Qiufeng Wang, Chang Liu, Xu Yang et al.AAAI 2026
- A Study on PAVE Specification for LearnwareHao-Yu Shi, Zhi-Hao Tan, Zi-Chen Zhao, Yang Yu et al.ICLR 2026
Builds on29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- EA-Vit: Efficient Adaptation for Elastic Vision TransformerChen Zhu, Wangbo Zhao, Huiwen Zhang, Yuhao Zhou et al.ICCV 2025 · 2 citations
- Slicing Vision Transformer for Flexible InferenceYitian Zhang, Huseyin Coskun, Xu Ma, Huan Wang et al.NeurIPS 2024 · 4 citations
- DeepCompress-ViT: Rethinking Model Compression to Enhance Efficiency of Vision Transformers at the EdgeSabbir Ahmed, Abdullah Al Arafat, Deniz Najafi, Akhlak Mahmood et al.CVPR 2025
- FlexRank: Nested Low-Rank Knowledge Decomposition for Adaptive Model DeploymentRiccardo Zaccone, Stefanos Laskaridis, Marco Ciccone, Samuel HorváthICML 2026 · 1 citation
- HydraViT: Stacking Heads for a Scalable ViTJanek Haberer, Ali Hojjat, Olaf LandsiedelNeurIPS 2024 · 11 citations
