AdaFuse: Adaptive Ensemble Decoding for Large Language Models
Chengming Cui, Tianxin Wei, Ziyi Chen, Ruizhong Qiu, Zhichen Zeng, Zhining Liu, Xuying Ning, Duo Zhou, Jingrui He
Abstract
Large language models (LLMs) exhibit complementary strengths arising from differences in pretraining data, model architectures, and decoding behaviors. Inference-time ensembling provides a practical way to combine these capabilities without retraining. However, existing ensemble approaches suffer from fundamental limitations. Most rely on fixed fusion granularity, which lacks the flexibility required for mid-generation adaptation and fails to adapt to different generation characteristics across tasks. To address these challenges, we propose ADAFUSE, an adaptive ensemble decoding framework that dynamically selects semantically appropriate fusion units during generation. Rather than committing to a fixed granularity, ADAFUSE adjusts fusion behavior on the fly based on the decoding context, with words serving as basic building blocks for alignment. To be specific, we introduce an uncertainty-based criterion to decide whether to apply ensembling at each decoding step. Under confident decoding states, the model continues generation directly. In less certain states, ADAFUSE invokes a diversity-aware scaling strategy to explore alternative candidate continuations and inform ensemble decisions. This design establishes a synergistic interaction between adaptive ensembling and test-time scaling, where ensemble decisions guide targeted exploration, and the resulting diversity in turn strengthens ensemble quality. Experiments on open-domain QA, arithmetic reasoning, and machine translation demonstrate that ADAFUSE consistently outperforms strong ensemble baselines, achieving an average relative improvement of 6.88%. The code is available at https://github.com/CCM0111/AdaFuse .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bea198bd-c3ba-4e88-8da1-2765656f2263Builds on12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionDongfu Jiang, Xiang Ren, Bill Yuchen LinACL 2023 · 95 citations
- Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel CollaborationYichong Huang, Xiaocheng Feng, Baohang Li, Yang Xiang et al.NeurIPS 2024 · 94 citations
Related papers
- Test-Time Scaling in Diffusion LLMS via Hidden Semi-Autoregressive ExpertsJihoon Lee, Hoyeon Moon, Kevin Zhai, Arun Kumar Chithanar et al.ICLR 2026 · 7 citations
- FUSE: Ensembling Verifiers with Zero Labeled DataJoonhyuk Lee, Virginia L., Sarah Zhao, Yash Nair et al.ICML 2026 · 2 citations
- SpecEM: Training-Free LLM Ensembling via Iterative Drafting, Verification, and Online FeedbackBo Lv, Nayu Liu, Chen Tang, Xin Liu et al.NeurIPS 2025 · 7 citations
- It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech RecognitionChen Chen, Ruizhe Li, Yuchen Hu, Sabato Marco Siniscalchi et al.ICLR 2024 · 37 citations
- Future-Gain Guided Test-Time Learning for Large Language ModelsLangYu Bian, Jinwu Hu, Zitian Zhang, Dongjin Yang et al.ICML 2026 · 12 citations
