Can We Scale Transformers to Predict Parameters of Diverse ImageNet Models?
Boris Knyazev, Doha Hwang, Simon Lacoste-Julien
Abstract
Pretraining a neural network on a large dataset is becoming a cornerstone in machine learning that is within the reach of only a few communities with large-resources. We aim at an ambitious goal of democratizing pretraining. Towards that goal, we train and release a single neural network that can predict high quality ImageNet parameters of other neural networks. By using predicted parameters for initialization we are able to boost training of diverse ImageNet models available in PyTorch. When transferred to other datasets, models initialized with predicted parameters also converge faster and reach competitive final performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a00bd7e-5967-4935-b0a6-6bf400c28467Cited by top-tier papers18
- Towards Scalable and Versatile Weight Space LearningKonstantin Schürholt, Michael W. Mahoney, Damian BorthICML 2024 · 39 citations
- Scale Equivariant Graph MetanetworksIoannis Kalogeropoulos, Giorgos Bouritsas, Yannis PanagakisNeurIPS 2024 · 24 citations
- Generative Modeling of Weights: Generalization or Memorization?Boya Zeng, Yida Yin, Zhiqiu Xu, Zhuang LiuCVPR 2026 · 12 citations
- Learning to Learn Weight Generation via Local Consistency DiffusionYunchuan Guan, Yu Liu, Ke Zhou, Zhiqi Shen et al.CVPR 2026 · 5 citations
- ECO: Evolving Core Knowledge for Efficient TransferFu Feng, Yucheng Xie, Ruixiao Shi, Jianlu Shen et al.NeurIPS 2025 · 4 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
Related papers
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 111 citations
- Mimetic Initialization of Self-Attention LayersAsher Trockman, J. Zico KolterICML 2023 · 54 citations
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou et al.AAAI 2020 · 219 citations
- ImageNet Pre-training Also Transfers Non-robustnessJiaming Zhang, Jitao Sang, Qi Yi, Yunfan Yang et al.AAAI 2023 · 6 citations
- Transferring Learning Trajectories of Neural NetworksDaiki ChijiwaICLR 2024 · 4 citations
