Equivariant Deep Weight Space Alignment
Aviv Navon, Aviv Shamsian, Ethan Fetaya, Gal Chechik, Nadav Dym, Haggai Maron
Abstract
Permutation symmetries of deep networks make basic operations like model merging and similarity estimation challenging. In many cases, aligning the weights of the networks, i.e., finding optimal permutations between their weights, is necessary. Unfortunately, weight alignment is an NP-hard problem. Prior research has mainly focused on solving relaxed versions of the alignment problem, leading to either time-consuming methods or sub-optimal solutions. To accelerate the alignment process and improve its quality, we propose a novel framework aimed at learning to solve the weight alignment problem, which we name Deep-Align. To that end, we first prove that weight alignment adheres to two fundamental symmetries and then, propose a deep architecture that respects these symmetries. Notably, our framework does not require any labeled data. We provide a theoretical analysis of our approach and evaluate Deep-Align on several types of network architectures and learning setups. Our experimental results indicate that a feed-forward pass with Deep-Align produces better or equivalent alignments compared to those produced by current optimization algorithms. Additionally, our alignments can be used as an effective initialization for other methods, leading to improved solutions with a significant speedup in convergence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d22363da-435b-4d4a-a119-1485277e5507Cited by top-tier papers19
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron et al.NeurIPS 2024 · 25 citations
- Scale Equivariant Graph MetanetworksIoannis Kalogeropoulos, Giorgos Bouritsas, Yannis PanagakisNeurIPS 2024 · 24 citations
- : Cycle-Consistent Multi-Model MergingDonato Crisostomi, Marco Fumero, Daniele Baieri, Florian Bernard et al.NeurIPS 2024 · 23 citations
- Improved Generalization of Weight Space Networks via AugmentationsAviv Shamsian, Aviv Navon, David W. Zhang, Yan Zhang et al.ICML 2024 · 19 citations
- GradMetaNet: An Equivariant Architecture for Learning on GradientsYoav Gelberg, Yam Eitan, Aviv Navon, Aviv Shamsian et al.NeurIPS 2025 · 8 citations
Builds on17
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
- Deep Graph Matching ConsensusMatthias Fey, Jan Eric Lenssen, Christopher Morris, Jonathan Masci et al.ICLR 2020 · 227 citations
- ZipIt! Merging Models from Different Tasks without TrainingGeorge Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh et al.ICLR 2024 · 185 citations
Related papers
- Align, then memorise: the dynamics of learning with feedback alignmentMaria Refinetti, Stéphane d'Ascoli, Ruben Ohana, Sebastian GoldtICML 2021 · 47 citations
- Deep Neural Network Fusion via Graph Matching with Applications to Model Ensemble and Federated LearningChang Liu, Chenfei Lou, Runzhong Wang, Alan Yuhan Xi et al.ICML 2022 · 72 citations
- Direct Feedback Alignment Scales to Modern Deep Learning Tasks and ArchitecturesJulien Launay, Iacopo Poli, François Boniface, Florent KrzakalaNeurIPS 2020 · 94 citations
- Optimizing Mode Connectivity via Neuron AlignmentN. Joseph Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk et al.NeurIPS 2020 · 104 citations
- Path-conditioned training: a principled way to rescale ReLU neural networksArthur Lebeurrier, Titouan Vayer, Rémi GribonvalICML 2026 · 3 citations
