Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights
Konstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian Borth
Abstract
Learning representations of neural network weights given a model zoo is an emerging and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we extend hyper-representations for generative use to sample new model weights. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and several sampling methods based on the topology of hyper-representations. The models generated using our methods are diverse, performant and capable to outperform strong baselines as evaluated on several downstream tasks: initialization, ensemble sampling and transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1e5d475-20cf-4845-8236-9e39208ba1edCited by top-tier papers38
- Equivariant Architectures for Learning in Deep Weight SpacesAviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya et al.ICML 2023 · 101 citations
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen et al.ICLR 2024 · 57 citations
- Towards Scalable and Versatile Weight Space LearningKonstantin Schürholt, Michael W. Mahoney, Damian BorthICML 2024 · 39 citations
- Sampling weights of deep neural networksErik Lien Bolager, Iryna Burak, Chinmay Datar, Qing Sun et al.NeurIPS 2023 · 36 citations
- Spatio-Temporal Few-Shot Learning via Diffusive Neural Network GenerationYuan Yuan, Chenyang Shao, Jingtao Ding, Depeng Jin et al.ICLR 2024 · 34 citations
Builds on9
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 1,188 citations
- A Baseline for Few-Shot Image ClassificationGuneet Singh Dhillon, Pratik Chaudhari, Avinash Ravichandran, Stefano SoattoICLR 2020 · 640 citations
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 467 citations
- From Variational to Deterministic AutoencodersPartha Ghosh, Mehdi S. M. Sajjadi, Antonio Vergari, Michael J. Black et al.ICLR 2020 · 298 citations
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 111 citations
Related papers
- Set-based Neural Network Encoding Without Weight TyingBruno Andreis, Bedionita Soro, Philip H. S. Torr, Sung Ju HwangNeurIPS 2024 · 8 citations
- Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic PredictionKonstantin Schürholt, Dimche Kostadinov, Damian BorthNeurIPS 2021 · 69 citations
- Hypernetwork approach to generating point cloudsPrzemyslaw Spurek, Sebastian Winczowski, Jacek Tabor, Maciej Zamorski et al.ICML 2020 · 36 citations
- Task-Adaptive Neural Network Search with Meta-Contrastive LearningWonyong Jeong, Hayeon Lee, Geon Park, Eunyoung Hyung et al.NeurIPS 2021 · 17 citations
- Generative Modeling of Weights: Generalization or Memorization?Boya Zeng, Yida Yin, Zhiqiu Xu, Zhuang LiuCVPR 2026 · 12 citations
