Making, Not Taking, the Best of N
Ammar Khairi, Daniel D'souza, Marzieh Fadaee, Julia Kreutzer
Abstract
Obtaining high-quality generations in modern LLMs has largely been framed as a selection problem: identifying a single winning generation from a diverse pool of samples, the Best-of- (BoN). Yet, this approach is inherently zero-sum, discarding diverse and potentially useful information from the pool. Instead, we explore a collaborative setup, where all candidates can potentially contribute to the final winning generation. To this end, we propose Fusion-of- (FusioN): a method that uses a general LLM judge to synthesize the most informative elements of each sample into a single final answer. We compare FusioN to BoN in two settings, (i) test-time scaling, where we sample and aggregate from a single model at test-time (ii) synthetic data generation, where we fuse samples from a pool of diverse teachers to improve a student model. We extensively benchmark both setups across 11 languages, 3 diverse benchmarks and varying model scales. Across the bench, FusionN consistently outperforms BoN showing versatility and robustness both in test-time scaling and in downstream gains from synthetic data generation. We also perform extensive analysis on FusioN, where it shows surprising strengths and robustness under challenging settings. These results show that we should shift how we think about evaluating and utilizing LLM generations from a monolithic measure of quality, to embracing their polylithic nature. This shift allows us to integrate diverse strengths, unlock latent potential, and achieve improvements that were previously inaccessible through selection alone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b31b47b-ced4-48b4-83c8-30b742d394b9Cited by top-tier papers1
Ask how each one uses itBuilds on26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
Related papers
- Don't Rank, Combine! Combining Machine Translation Hypotheses Using Quality EstimationGiorgos Vernikos, Andrei Popescu-BelisACL 2024
- FUSE: Ensembling Verifiers with Zero Labeled DataJoonhyuk Lee, Virginia L., Sarah Zhao, Yash Nair et al.ICML 2026 · 2 citations
- FuseGen: PLM Fusion for Data-generation based Zero-shot LearningTianyuan Zou, Yang Liu, Peng Li, Jianqing Zhang et al.EMNLP 2024 · 3 citations
- LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative FusionDongfu Jiang, Xiang Ren, Bill Yuchen LinACL 2023 · 95 citations
- AdaFuse: Adaptive Ensemble Decoding for Large Language ModelsChengming Cui, Tianxin Wei, Ziyi Chen, Ruizhong Qiu et al.ACL 2026
