Unlocking Pre-trained Weights: Parameter Inheritance for Zero-Shot Initialization
Jiaze Xu, Shiyu Xia, Jiaqi Lv, Xin Geng
Abstract
Appropriate parameter initialization is crucial for reducing the training cost of deep neural networks. Graph HyperNetworks (GHN) have emerged as a promising approach for initializing diverse architectures, with recent methods such as Task-Aware Learngene (TAL) further attempting to leverage pre-trained model knowledge via soft label supervision. However, such indirect supervision fails to fully exploit the rich information encoded in pre-trained weights. We propose Parameter InheriTance HyperNetwork (PITH), which introduces a novel parameter projection mechanism to directly inherit parameters from pretrained models for initializing target networks of varying configurations. Our method enables initialized networks to directly achieve competitive performance on downstream tasks without any further training, which we term zero-shot initialization. Extensive experiments demonstrate the superiority of PITH: ViT-Base initialized by PITH achieves 53.35% zero-shot accuracy on ImageNet-1K, surpassing the previous state-of-the-art by 6.54%, with consistent improvements across multiple downstream tasks. We provide the code at https://github.com/mathieuxu/PITH-Parameter-InheriTance-HyperNetwork.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae919c8c-34ee-4dba-a99a-2c0b2c502138Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 2,162 citations
- Knowledge distillation: A good teacher is patient and consistentLucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva et al.CVPR 2022 · 215 citations
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 111 citations
Related papers
- Learngene Tells You How to Customize: Task-Aware Parameter Initialization at Flexible ScalesJiaze Xu, Shiyu Xia, Xu Yang, Jiaqi Lv et al.ICML 2025
- Attribute Propagation Network for Graph Zero-Shot LearningLu Liu, Tianyi Zhou, Guodong Long, Jing Jiang et al.AAAI 2020 · 85 citations
- Adaptive-Learngene: Continual Expansion and Task-Aware Selection of Learngenes for Dynamic EnvironmentsShuxia Lin, Qiufeng Wang, Chang Liu, Xu Yang et al.AAAI 2026
- Towards Graph Foundation Models: Learning Generalities Across Graphs via Task-TreesZehong Wang, Zheyuan Zhang, Tianyi Ma, Nitesh V. Chawla et al.ICML 2025
- Deep Model ReassemblyXingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye et al.NeurIPS 2022 · 162 citations
