Transport and Merge: Cross-Architecture Merging for Large Language Models
Chenhang Cui, Binyun Yang, Fei Shen, Yuxin Chen, Jingnan Zheng, Xiang Wang, An Zhang, Tat-Seng Chua
Abstract
Large language models (LLMs) achieve strong capabilities by scaling model capacity and training data, yet many real-world deployments rely on smaller models trained or adapted from low-resource data. This gap motivates the need for mechanisms to transfer knowledge from large, high-resource models to smaller, resource-constrained targets. While model merging provides an effective transfer mechanism, most existing approaches assume architecture-compatible models and therefore cannot directly transfer knowledge from large high-resource LLMs to heterogeneous lowresource targets. In this work, we propose a cross-architecture merging framework based on optimal transport (OT) that aligns activations to infer cross-neuron correspondences between heterogeneous models. The resulting transport matrices are then used to guide direct weightspace fusion, enabling effective high-resource to low-resource transfer using only a small set of inputs. Extensive experiments across low-resource languages and specialized domains demonstrate consistent improvements over target base models. The code is available at https: //github.com/chenhangcuisg-code/ Cross-Architecture-Merging-for- Large-Language-Models/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5b5bde5-1cf5-47e4-a8d0-826f95b1771bBuilds on18
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 330 citations
- Task Arithmetic in the Tangent Space: Improved Editing of Pre-Trained ModelsGuillermo Ortiz-Jiménez, Alessandro Favero, Pascal FrossardNeurIPS 2023 · 272 citations
Related papers
- Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible MultilingualityMengyu Bu, Yang FengACL 2026 · 2 citations
- LS-Merge: Merging Language Models in Latent SpaceBedionita Soro, Aoxuan Silvia Zhang, Bruno Andreis, Jaehyeong Jo et al.ICLR 2026
- Knowledge Fusion of Large Language ModelsFanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan et al.ICLR 2024 · 113 citations
- Probabilistic Token Alignment for Large Language Model FusionRunjia Zeng, James Liang, Cheng Han, Zhiwen Cao et al.NeurIPS 2025 · 3 citations
- MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language ModelsHoang Tran Vuong, Tue Le, Quyen Tran, Linh Ngo Van et al.AAAI 2026
