Cool-Fusion: Fuse Large Language Models without Training
Cong Liu, Xiaojun Quan, Yan Pan, Weigang Wu, Xu Chen, Liang Lin
Abstract
We focus on the problem of fusing two or more heterogeneous large language models (LLMs) to leverage their complementary strengths. One of the challenges of model fusion is high computational load, specifically in fine-tuning or aligning vocabularies. To address this, we propose Cool-Fusion, a simple yet effective approach that fuses the knowledge of source LLMs, which does not require training. Unlike ensemble methods, Cool-Fusion is applicable to any set of source LLMs that have different vocabularies. To overcome the vocabulary discrepancies among LLMs, we ensemble LLMs on text level, allowing them to rerank the generated texts by each other with different granularities. Extensive experiments have been conducted across a variety of benchmark datasets. On GSM8K, Cool-Fusion increases accuracy from three strong source LLMs by a significant margin of 17.4%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fc0d581-fb1a-4b55-84dd-be1a9ed0738aCited by top-tier papers6
- Sampling from Your Language Model One Byte at a TimeJonathan Hayase, Alisa Liu, Noah Smith, Sewoong OhICML 2026 · 9 citations
- When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM EnsemblingHeecheol Yun, Kwangmin Ki, Jung Hyun Lee, Eunho YangICLR 2026 · 3 citations
- Lossless Vocabulary Reduction for Auto-Regressive Language ModelsDaiki Chijiwa, Taku Hasegawa, Kyosuke Nishida, Shin'ya Yamaguchi et al.ICLR 2026 · 3 citations
- Collaborative Beam Search: Enhancing LLM Reasoning via Collective ConsensusYangyifan Xu, Shuo Ren, Jiajun ZhangEMNLP 2025 · 1 citation
- Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model EnsemblingYuxuan Yao, Han Wu, Mingyang Liu, Sichun Luo et al.ICLR 2025
Builds on12
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Specializing Smaller Language Models towards Multi-Step ReasoningYao Fu, Hao Peng, Litu Ou, Ashish Sabharwal et al.ICML 2023 · 347 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- Scalable Best-of-N Selection for Large Language Models via Self-CertaintyZhewei Kang, Xuandong Zhao, Dawn SongNeurIPS 2025 · 211 citations
- SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive SummarizationMathieu Ravaut, Shafiq R. Joty, Nancy F. ChenACL 2022 · 116 citations
Related papers
- Knowledge Fusion of Large Language ModelsFanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan et al.ICLR 2024 · 113 citations
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang et al.NeurIPS 2024 · 94 citations
- Probabilistic Token Alignment for Large Language Model FusionRunjia Zeng, James Liang, Cheng Han, Zhiwen Cao et al.NeurIPS 2025 · 3 citations
- FuseChat: Knowledge Fusion of Chat ModelsFanqi Wan, Longguang Zhong, Ziyi Yang, Ruijun Chen et al.EMNLP 2025 · 4 citations
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionDongfu Jiang, Xiang Ren, Bill Yuchen LinACL 2023 · 95 citations
