Transformer Fusion with Optimal Transport
Moritz Imfeld, Jacopo Graldi, Marco Giordano, Thomas Hofmann, Sotiris Anagnostidis, Sidak Pal Singh
摘要
Fusion is a technique for merging multiple independently-trained neural networks in order to combine their capabilities. Past attempts have been restricted to the case of fully-connected, convolutional, and residual networks. This paper presents a systematic approach for fusing two or more transformer-based networks exploiting Optimal Transport to (soft-)align the various architectural components. We flesh out an abstraction for layer alignment, that can generalize to arbitrary architectures - in principle - and we apply this to the key ingredients of Transformers such as multi-head self-attention, layer-normalization, and residual connections, and we discuss how to handle them via various ablation studies. Furthermore, our method allows the fusion of models of different sizes (heterogeneous fusion), providing a new and efficient way to compress Transformers. The proposed approach is evaluated on both image classification tasks via Vision Transformer and natural language modeling tasks using BERT. Our approach consistently outperforms vanilla fusion, and, after a surprisingly short finetuning, also outperforms the individual converged parent models. In our analysis, we uncover intriguing insights about the significant role of soft alignment in the case of Transformers. Our results showcase the potential of fusing multiple Transformers, thus compounding their expertise, in the budding paradigm of model fusion and recombination. Code is available at https://github.com/graldij/transformer-fusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Model Fusion through Bayesian Optimization in Language Model Fine-TuningChaeyun Jang, Hyungi Lee, Jungtaek Kim, Juho LeeNeurIPS 2024 · 被引用 8 次
- Transporting Task Vectors across Different Architectures without TrainingFilippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello 等ICML 2026 · 被引用 3 次
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 被引用 1 次
- Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-ExpertsQi Wang, Hanyang Peng, Yue YuAAAI 2026 · 被引用 1 次
- Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model FusionBinchi Zhang, Zaiyi Zheng, Zhengzhang Chen, Jundong LiICML 2025
它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos 等ICLR 2020 · 被引用 1,368 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
相关 Paper
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 被引用 330 次
- Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal TransportShaan Shah, Meenakshi KhoslaICLR 2026 · 被引用 4 次
- LS-Merge: Merging Language Models in Latent SpaceBedionita Soro, Aoxuan Silvia Zhang, Bruno Andreis, Jaehyeong Jo 等ICLR 2026
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang 等CVPR 2022 · 被引用 214 次
- TaskFusion: An Efficient Transfer Learning Architecture with Dual Delta Sparsity for Multi-Task Natural Language ProcessingZichen Fan, Qirui Zhang, Pierre Abillama, Sara Shoouri 等ISCA 2023 · 被引用 15 次
