Universal Neural Functionals
Allan Zhou, Chelsea Finn, James Harrison
Abstract
A challenging problem in many modern machine learning tasks is to process weight-space features, i.e., to transform or extract information from the weights and gradients of a neural network. Recent works have developed promising weight-space models that are equivariant to the permutation symmetries of simple feedforward networks. However, they are not applicable to general architectures, since the permutation symmetries of a weight space can be complicated by recurrence or residual connections. This work proposes an algorithm that automatically constructs permutation equivariant models, which we refer to as universal neural functionals (UNFs), for any weight space. Among other applications, we demonstrate how UNFs can be substituted into existing learned optimizer designs, and find promising improvements over prior methods when optimizing small image classifiers and language models. Our results suggest that learned optimizers can benefit from considering the (symmetry) structure of the weight space they optimize. We open-source our library for constructing UNFs at https://github.com/AllanYangZhou/universal_neural_functional.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron et al.NeurIPS 2024 · 25 citations
- Scale Equivariant Graph MetanetworksIoannis Kalogeropoulos, Giorgos Bouritsas, Yannis PanagakisNeurIPS 2024 · 24 citations
- Monomial Matrix Group Equivariant Neural Functional NetworksHoang V. Tran, Thieu N. Vo, Tho Huu, An Nguyen The et al.NeurIPS 2024 · 19 citations
- LLaNA: Large Language and NeRF AssistantAndrea Amaduzzi, Pierluigi Zama Ramirez, Giuseppe Lisanti, Samuele Salti et al.NeurIPS 2024 · 10 citations
- GradMetaNet: An Equivariant Architecture for Learning on GradientsYoav Gelberg, Yam Eitan, Aviv Navon, Aviv Shamsian et al.NeurIPS 2025 · 8 citations
Builds on12
- A Practical Method for Constructing Equivariant Multilayer Perceptrons for Arbitrary Matrix GroupsMarc Finzi, Max Welling, Andrew Gordon WilsonICML 2021 · 226 citations
- Equivariant Architectures for Learning in Deep Weight SpacesAviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya et al.ICML 2023 · 101 citations
- Permutation Equivariant Neural FunctionalsAllan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace et al.NeurIPS 2023 · 84 citations
- On the Symmetries of Deep Learning Models and their Internal RepresentationsCharles Godfrey, Davis Brown, Tegan Emerson, Henry KvingeNeurIPS 2022 · 78 citations
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 77 citations
Related papers
- Neural Functional TransformersAllan Zhou, Kaien Yang, Yiding Jiang, Kaylee Burns et al.NeurIPS 2023 · 53 citations
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen et al.ICLR 2024 · 57 citations
- On the Expressive Power of Permutation-Equivariant Weight-Space NetworksAdir Dayan, Yam Eitan, Haggai MaronICML 2026
- Equivariant Neural Functional Networks for TransformersHoang V. Tran, Thieu Vo, An Nguyen The, Tho Tran Huu et al.ICLR 2025
- Equivariant Machine Learning on Graphs with Nonlinear Spectral FiltersYa-Wei Eileen Lin, Ronen Talmon, Ron LevieNeurIPS 2024 · 5 citations
