Principled Weight Initialization for Hypernetworks
Oscar Chang, Lampros Flokas, Hod Lipson
摘要
Hypernetworks are meta neural networks that generate weights for a main neural network in an end-to-end differentiable manner. Despite extensive applications ranging from multi-task learning to Bayesian deep learning, the problem of optimizing hypernetworks has not been studied to date. We observe that classical weight initialization methods like Glorot & Bengio (2010) and He et al. ( 2015 ), when applied directly on a hypernet, fail to produce weights for the mainnet in the correct scale. We develop principled techniques for weight initialization in hypernets, and show that they lead to more stable mainnet weights, lower training loss, and faster convergence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 被引用 111 次
- Equivariant Architectures for Learning in Deep Weight SpacesAviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya 等ICML 2023 · 被引用 101 次
- On the Modularity of HypernetworksTomer Galanti, Lior WolfNeurIPS 2020 · 被引用 79 次
- Posterior Meta-Replay for Continual LearningChristian Henning, Maria R. Cervera, Francesco D'Angelo, Johannes von Oswald 等NeurIPS 2021 · 被引用 78 次
- DUET: A Tuning-Free Device-Cloud Collaborative Parameters Generation Framework for Efficient Device Model GeneralizationZheqi Lv, Wenqiao Zhang, Shengyu Zhang, Kun Kuang 等WWW 2023 · 被引用 68 次
它引用的顶会 Paper1
相关 Paper
- Precise characterization of the prior predictive distribution of deep ReLU networksLorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin 等NeurIPS 2021 · 被引用 36 次
- Magnitude Invariant Parametrizations Improve Hypernetwork LearningJose Javier Gonzalez Ortiz, John V. Guttag, Adrian V. DalcaICLR 2024 · 被引用 13 次
- Amortising Inference and Meta-Learning Priors in Neural NetworksTommy Rochussen, Vincent FortuinICLR 2026
- HyperNet Fields: Efficiently Training Hypernetworks without Ground Truth by Learning Weight TrajectoriesEric Hedlin, Munawar Hayat, Fatih Porikli, Kwang Moo Yi 等CVPR 2025
- Hyper-Representations as Generative Models: Sampling Unseen Neural Network WeightsKonstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian BorthNeurIPS 2022 · 被引用 78 次
