Deep Linear Probe Generators for Weight Space Learning
Jonathan Kahana, Eliahu Horwitz, Imri Shuval, Yedid Hoshen
摘要
Weight space learning aims to extract information about a neural network, such as its training dataset or generalization error. Recent approaches learn directly from model weights, but this presents many challenges as weights are high-dimensional and include permutation symmetries between neurons. An alternative approach, Probing, represents a model by passing a set of learned inputs (probes) through the model, and training a predictor on top of the corresponding outputs. Although probing is typically not used as a stand alone approach, our preliminary experiment found that a vanilla probing baseline worked surprisingly well. However, we discover that current probe learning strategies are ineffective. We therefore propose Deep Linear Probe Generators (ProbeGen), a simple and effective modification to probing approaches. ProbeGen adds a shared generator module with a deep linear architecture, providing an inductive bias towards structured probes thus reducing overfitting. While simple, ProbeGen performs significantly better than the state-of-the-art and is very efficient, requiring between 30 to 1000 times fewer FLOPs than other top approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- GradMetaNet: An Equivariant Architecture for Learning on GradientsYoav Gelberg, Yam Eitan, Aviv Navon, Aviv Shamsian 等NeurIPS 2025 · 被引用 8 次
- Weight-Space Linear Recurrent Neural NetworksRoussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo-Fernandez, Idriss Tsayem 等ICLR 2026 · 被引用 6 次
- What Linear Probes Miss: Multi-View Probing for Weight-Space LearningEunwoo Heo, Kyeongkook Seo, Jaejun YooICML 2026
- Parameter Manifold PurificationJiacong Hu, Jinxun Wu, Shengxuming Zhang, Shunyu Liu 等ICML 2026
- Learning on Model Weights using Tree ExpertsEliahu Horwitz, Bar Cavia, Jonathan Kahana, Yedid HoshenCVPR 2025
它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell 等NeurIPS 2020 · 被引用 4,008 次
- SOK: (State of) The Art of War: Offensive Techniques in Binary AnalysisYan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens 等S&P 2016 · 被引用 1,085 次
相关 Paper
- Learning Useful Representations of Recurrent Neural Network Weight MatricesVincent Herrmann, Francesco Faccio, Jürgen SchmidhuberICML 2024 · 被引用 12 次
- On the Expressive Power of Permutation-Equivariant Weight-Space NetworksAdir Dayan, Yam Eitan, Haggai MaronICML 2026
- Towards Scalable and Versatile Weight Space LearningKonstantin Schürholt, Michael W. Mahoney, Damian BorthICML 2024 · 被引用 39 次
- Set-based Neural Network Encoding Without Weight TyingBruno Andreis, Bedionita Soro, Philip H. S. Torr, Sung Ju HwangNeurIPS 2024 · 被引用 8 次
- A Graph Meta-Network for Learning on Kolmogorov–Arnold NetworksGuy Bar-Shalom, Ami Tavory, Itay Evron, Maya Bechler-Speicher 等ICLR 2026 · 被引用 1 次
