Towards Scalable and Versatile Weight Space Learning
Konstantin Schürholt, Michael W. Mahoney, Damian Borth
摘要
Learning representations of well-trained neural network models holds the promise to provide an understanding of the inner workings of those models. However, previous work has either faced limitations when processing larger networks or was task-specific to either discriminative or generative tasks. This paper introduces the SANE approach to weight-space learning. SANE overcomes previous limitations by learning task-agnostic representations of neural networks that are scalable to larger models of varying architectures and that show capabilities beyond a single task. Our method extends the idea of hyper-representations towards sequential processing of subsets of neural network weights, thus allowing one to embed larger neural networks as a set of tokens into the learned representation space. SANE reveals global model information from layer-wise embeddings, and it can sequentially generate unseen neural network models, which was unattainable with previous hyper-representation learning methods. Extensive empirical evaluation demonstrates that SANE matches or exceeds state-of-the-art performance on several weight representation learning benchmarks, particularly in initialization for new tasks and larger ResNet architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsZhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao 等NeurIPS 2025 · 被引用 22 次
- Generative Modeling of Weights: Generalization or Memorization?Boya Zeng, Yida Yin, Zhiqiu Xu, Zhuang LiuCVPR 2026 · 被引用 12 次
- Scaling Up Parameter Generation: A Recurrent Diffusion ApproachKai Wang, Dongwen Tang, Wangbo Zhao, Konstantin Schürholt 等NeurIPS 2025 · 被引用 9 次
- GradMetaNet: An Equivariant Architecture for Learning on GradientsYoav Gelberg, Yam Eitan, Aviv Navon, Aviv Shamsian 等NeurIPS 2025 · 被引用 8 次
- Weight-Space Linear Recurrent Neural NetworksRoussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo-Fernandez, Idriss Tsayem 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Parameter Prediction for Unseen Deep ArchitecturesBoris Knyazev, Michal Drozdzal, Graham W. Taylor, Adriana Romero-SorianoNeurIPS 2021 · 被引用 111 次
- Learning Neural Network SubspacesMitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi 等ICML 2021 · 被引用 101 次
- Loss Surface Simplexes for Mode Connecting Volumes and Fast EnsemblingGregory W. Benton, Wesley J. Maddox, Sanae Lotfi, Andrew Gordon WilsonICML 2021 · 被引用 88 次
相关 Paper
- Hyper-Representations as Generative Models: Sampling Unseen Neural Network WeightsKonstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian BorthNeurIPS 2022 · 被引用 78 次
- Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic PredictionKonstantin Schürholt, Dimche Kostadinov, Damian BorthNeurIPS 2021 · 被引用 69 次
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 被引用 412 次
- Learning Useful Representations of Recurrent Neural Network Weight MatricesVincent Herrmann, Francesco Faccio, Jürgen SchmidhuberICML 2024 · 被引用 12 次
- Deep Linear Probe Generators for Weight Space LearningJonathan Kahana, Eliahu Horwitz, Imri Shuval, Yedid HoshenICLR 2025
