Transformer Fusion with Optimal Transport
Moritz Imfeld, Jacopo Graldi, Marco Giordano, Thomas Hofmann, Sotiris Anagnostidis, Sidak Pal Singh
Abstract
Fusion is a technique for merging multiple independently-trained neural networks in order to combine their capabilities. Past attempts have been restricted to the case of fully-connected, convolutional, and residual networks. This paper presents a systematic approach for fusing two or more transformer-based networks exploiting Optimal Transport to (soft-)align the various architectural components. We flesh out an abstraction for layer alignment, that can generalize to arbitrary architectures - in principle - and we apply this to the key ingredients of Transformers such as multi-head self-attention, layer-normalization, and residual connections, and we discuss how to handle them via various ablation studies. Furthermore, our method allows the fusion of models of different sizes (heterogeneous fusion), providing a new and efficient way to compress Transformers. The proposed approach is evaluated on both image classification tasks via Vision Transformer and natural language modeling tasks using BERT. Our approach consistently outperforms vanilla fusion, and, after a surprisingly short finetuning, also outperforms the individual converged parent models. In our analysis, we uncover intriguing insights about the significant role of soft alignment in the case of Transformers. Our results showcase the potential of fusing multiple Transformers, thus compounding their expertise, in the budding paradigm of model fusion and recombination. Code is available at https://github.com/graldij/transformer-fusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7eef5dd-438b-4e3f-8f09-37b1b8083231Cited by top-tier papers10
- Model Fusion through Bayesian Optimization in Language Model Fine-TuningChaeyun Jang, Hyungi Lee, Jungtaek Kim, Juho LeeNeurIPS 2024 · 8 citations
- Transporting Task Vectors across Different Architectures without TrainingFilippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello et al.ICML 2026 · 3 citations
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 1 citation
- Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-ExpertsQi Wang, Hanyang Peng, Yue YuAAAI 2026 · 1 citation
- Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model FusionBinchi Zhang, Zaiyi Zheng, Zhengzhang Chen, Jundong LiICML 2025
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 330 citations
- Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal TransportShaan Shah, Meenakshi KhoslaICLR 2026 · 4 citations
- LS-Merge: Merging Language Models in Latent SpaceBedionita Soro, Aoxuan Silvia Zhang, Bruno Andreis, Jaehyeong Jo et al.ICLR 2026
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
- TaskFusion: An Efficient Transfer Learning Architecture with Dual Delta Sparsity for Multi-Task Natural Language ProcessingZichen Fan, Qirui Zhang, Pierre Abillama, Sara Shoouri et al.ISCA 2023 · 15 citations
