Knowledge Fusion of Large Language Models
Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, Shuming Shi
Abstract
While training large language models (LLMs) from scratch can generate models with distinct functionalities and strengths, it comes at significant costs and may result in redundant capabilities. Alternatively, a cost-effective and compelling approach is to merge existing pre-trained LLMs into a more potent model. However, due to the varying architectures of these LLMs, directly blending their weights is impractical. In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM. By leveraging the generative distributions of source LLMs, we externalize their collective knowledge and unique strengths, thereby potentially elevating the capabilities of the target model beyond those of any individual source LLM. We validate our approach using three popular LLMs with different architectures--Llama-2, MPT, and OpenLLaMA--across various benchmarks and tasks. Our findings confirm that the fusion of LLMs can improve the performance of the target model across a range of capabilities such as reasoning, commonsense, and code generation. Our code, model weights, and data are public at https://github.com/fanqiwan/FuseLLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1d9bf81-95e6-4ffe-99eb-9f7754a7c476Cited by top-tier papers73
- Reversing the Forget-Retain Objectives: An Efficient LLM Unlearning Framework from Logit DifferenceJiabao Ji, Yujian Liu, Yang Zhang, Gaowen Liu et al.NeurIPS 2024 · 106 citations
- Not All Tokens Are What You Need for PretrainingZhenghao Lin, Zhibin Gou, Yeyun Gong, Xiao Liu et al.NeurIPS 2024 · 99 citations
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang et al.NeurIPS 2024 · 94 citations
- Parameter Competition Balancing for Model MergingGuodong Du, Junlin Lee, Jing Li, Runhua Jiang et al.NeurIPS 2024 · 91 citations
- Universal Cross-Tokenizer Distillation via Approximate Likelihood MatchingBenjamin Minixhofer, Ivan Vulic, Edoardo Maria PontiNeurIPS 2025 · 48 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho et al.NeurIPS 2021 · 630 citations
Related papers
- FuseChat: Knowledge Fusion of Chat ModelsFanqi Wan, Longguang Zhong, Ziyi Yang, Ruijun Chen et al.EMNLP 2025 · 4 citations
- Probabilistic Token Alignment for Large Language Model FusionRunjia Zeng, James Liang, Cheng Han, Zhiwen Cao et al.NeurIPS 2025 · 3 citations
- Weighted-Reward Preference Optimization for Implicit Model FusionZiyi Yang, Fanqi Wan, Longguang Zhong, Tianyuan Shi et al.ICLR 2025
- Cool-Fusion: Fuse Large Language Models without TrainingCong Liu, Xiaojun Quan, Yan Pan, Weigang Wu et al.ACL 2025 · 12 citations
- LMFusion: Adapting Pretrained Language Models for Multimodal GenerationWeijia Shi, Xiaochuang Han, Chunting Zhou, Weixin Liang et al.NeurIPS 2025 · 134 citations
