Principled Weight Initialization for Hypernetworks
Oscar Chang, Lampros Flokas, Hod Lipson
Abstract
Hypernetworks are meta neural networks that generate weights for a main neural network in an end-to-end differentiable manner. Despite extensive applications ranging from multi-task learning to Bayesian deep learning, the problem of optimizing hypernetworks has not been studied to date. We observe that classical weight initialization methods like Glorot & Bengio (2010) and He et al. ( 2015 ), when applied directly on a hypernet, fail to produce weights for the mainnet in the correct scale. We develop principled techniques for weight initialization in hypernets, and show that they lead to more stable mainnet weights, lower training loss, and faster convergence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers28
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 111 citations
- Equivariant Architectures for Learning in Deep Weight SpacesAviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya et al.ICML 2023 · 101 citations
- On the Modularity of HypernetworksTomer Galanti, Lior WolfNeurIPS 2020 · 79 citations
- Posterior Meta-Replay for Continual LearningChristian Henning, Maria R. Cervera, Francesco D'Angelo, Johannes von Oswald et al.NeurIPS 2021 · 78 citations
- DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model GeneralizationZheqi Lv, Wenqiao Zhang, Shengyu Zhang, Kun Kuang et al.WWW 2023 · 68 citations
Builds on1
Related papers
- Precise characterization of the prior predictive distribution of deep ReLU networksLorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin et al.NeurIPS 2021 · 36 citations
- Magnitude Invariant Parametrizations Improve Hypernetwork LearningJose Javier Gonzalez Ortiz, John V. Guttag, Adrian V. DalcaICLR 2024 · 13 citations
- Amortising Inference and Meta-Learning Priors in Neural NetworksTommy Rochussen, Vincent FortuinICLR 2026
- HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight TrajectoriesEric Hedlin, Munawar Hayat, Fatih Porikli, Kwang Moo Yi et al.CVPR 2025
- Hyper-Representations as Generative Models: Sampling Unseen Neural Network WeightsKonstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian BorthNeurIPS 2022 · 78 citations
