Equivariant Deep Weight Space Alignment
Aviv Navon, Aviv Shamsian, Ethan Fetaya, Gal Chechik, Nadav Dym, Haggai Maron
摘要
Permutation symmetries of deep networks make basic operations like model merging and similarity estimation challenging. In many cases, aligning the weights of the networks, i.e., finding optimal permutations between their weights, is necessary. Unfortunately, weight alignment is an NP-hard problem. Prior research has mainly focused on solving relaxed versions of the alignment problem, leading to either time-consuming methods or sub-optimal solutions. To accelerate the alignment process and improve its quality, we propose a novel framework aimed at learning to solve the weight alignment problem, which we name Deep-Align. To that end, we first prove that weight alignment adheres to two fundamental symmetries and then, propose a deep architecture that respects these symmetries. Notably, our framework does not require any labeled data. We provide a theoretical analysis of our approach and evaluate Deep-Align on several types of network architectures and learning setups. Our experimental results indicate that a feed-forward pass with Deep-Align produces better or equivalent alignments compared to those produced by current optimization algorithms. Additionally, our alignments can be used as an effective initialization for other methods, leading to improved solutions with a significant speedup in convergence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron 等NeurIPS 2024 · 被引用 25 次
- Scale Equivariant Graph MetanetworksIoannis Kalogeropoulos, Giorgos Bouritsas, Yannis PanagakisNeurIPS 2024 · 被引用 24 次
- : Cycle-Consistent Multi-Model MergingDonato Crisostomi, Marco Fumero, Daniele Baieri, Florian Bernard 等NeurIPS 2024 · 被引用 23 次
- Improved Generalization of Weight Space Networks via AugmentationsAviv Shamsian, Aviv Navon, David W. Zhang, Yan Zhang 等ICML 2024 · 被引用 19 次
- GradMetaNet: An Equivariant Architecture for Learning on GradientsYoav Gelberg, Yam Eitan, Aviv Navon, Aviv Shamsian 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper17
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos 等ICLR 2020 · 被引用 1,368 次
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 被引用 301 次
- Deep Graph Matching ConsensusMatthias Fey, Jan Eric Lenssen, Christopher Morris, Jonathan Masci 等ICLR 2020 · 被引用 227 次
- ZipIt! Merging Models from Different Tasks without TrainingGeorge Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh 等ICLR 2024 · 被引用 185 次
相关 Paper
- Align, then memorise: the dynamics of learning with feedback alignmentMaria Refinetti, Stéphane d'Ascoli, Ruben Ohana, Sebastian GoldtICML 2021 · 被引用 47 次
- Deep Neural Network Fusion via Graph Matching with Applications to Model Ensemble and Federated LearningChang Liu, Chenfei Lou, Runzhong Wang, Alan Yuhan Xi 等ICML 2022 · 被引用 72 次
- Direct Feedback Alignment Scales to Modern Deep Learning Tasks and ArchitecturesJulien Launay, Iacopo Poli, François Boniface, Florent KrzakalaNeurIPS 2020 · 被引用 94 次
- Optimizing Mode Connectivity via Neuron AlignmentN. Joseph Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk 等NeurIPS 2020 · 被引用 104 次
- Path-conditioned training: a principled way to rescale ReLU neural networksArthur Lebeurrier, Titouan Vayer, Rémi GribonvalICML 2026 · 被引用 3 次
