Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction
Konstantin Schürholt, Dimche Kostadinov, Damian Borth
Abstract
Self-Supervised Learning (SSL) has been shown to learn useful and informationpreserving representations. Neural Networks (NNs) are widely applied, yet their weight space is still not fully understood. Therefore, we propose to use SSL to learn hyper-representations of the weights of populations of NNs. To that end, we introduce domain specific data augmentations and an adapted attention architecture. Our empirical evaluation demonstrates that self-supervised representation learning in this domain is able to recover diverse NN model characteristics. Further, we show that the proposed learned representations outperform prior work for predicting hyper-parameters, test accuracy, and generalization gap as well as transfer to out-of-distribution settings. Code and datasets are publicly available 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8f8e969-3a83-4732-955f-e978f6d2bf73Cited by top-tier papers36
- From data to functa: Your data point is a function and you can treat it like oneEmilien Dupont, Hyunjik Kim, S. M. Ali Eslami, Danilo Jimenez Rezende et al.ICML 2022 · 209 citations
- Equivariant Architectures for Learning in Deep Weight SpacesAviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya et al.ICML 2023 · 101 citations
- Hyper-Representations as Generative Models: Sampling Unseen Neural Network WeightsKonstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian BorthNeurIPS 2022 · 78 citations
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen et al.ICLR 2024 · 57 citations
- Interpreting the Weight Space of Customized Diffusion ModelsAmil Dravid, Yossi Gandelsman, Kuan-Chieh Wang, Rameen Abdal et al.NeurIPS 2024 · 40 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- Intriguing Properties of Contrastive LossesTing Chen, Calvin Luo, Lala LiNeurIPS 2021 · 206 citations
- Computing the Testing Error Without a Testing SetCiprian A. Corneanu, Sergio Escalera, Aleix M. MartinezCVPR 2020
Related papers
- Towards Scalable and Versatile Weight Space LearningKonstantin Schürholt, Michael W. Mahoney, Damian BorthICML 2024 · 39 citations
- Improved Generalization of Weight Space Networks via AugmentationsAviv Shamsian, Aviv Navon, David W. Zhang, Yan Zhang et al.ICML 2024 · 19 citations
- Reverse Engineering Self-Supervised LearningIdo Ben-Shaul, Ravid Shwartz-Ziv, Tomer Galanti, Shai Dekel et al.NeurIPS 2023 · 55 citations
- RSA: Reducing Semantic Shift from Aggressive Augmentations for Self-supervised LearningYingbin Bai, Erkun Yang, Zhaoqing Wang, Yuxuan Du et al.NeurIPS 2022 · 18 citations
- HypeBoy: Generative Self-Supervised Representation Learning on HypergraphsSunwoo Kim, Shinhwan Kang, Fanchen Bu, Soo Yong Lee et al.ICLR 2024 · 22 citations
