Best-of-Infinity: Asymptotic Performance of Test-Time LLM Ensembling
Junpei Komiyama, Daisuke Oba, Masafumi Oyamada
摘要
We study best-of- for large language models (LLMs) where the selection is based on majority voting. In particular, we analyze the limit , which we denote as best-of-. While this approach achieves impressive performance in the limit, it requires an infinite test-time budget. To address this, we propose an adaptive generation scheme that selects based on answer agreement, thereby efficiently allocating inference-time computation. Beyond adaptivity, we extend the framework to weighted ensembles of multiple LLMs, showing that such mixtures can outperform any individual model. The optimal ensemble weighting is formulated and efficiently computed as a mixed-integer linear program. Extensive experiments demonstrate the effectiveness of our approach. Our code is available at https://github.com/jkomiyama/BoInf-code-publish/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought ReasoningRenos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille 等ICML 2026 · 被引用 4 次
- Aligning Tree-Search Policies with Fixed Token Budgets in Test-Time Scaling of LLMsSora Miyamoto, Daisuke Oba, Naoaki OkazakiICML 2026 · 被引用 3 次
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for LLM Problem-SolvingYangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck 等ICLR 2025
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference ScalingQiwei Di, Kaixuan Ji, Xuheng Li, Heyang Zhao 等ICLR 2026 · 被引用 7 次
- Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-MakingLarkin Liu, Jalal EtesamiNeurIPS 2025 · 被引用 2 次
- Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability TheoryYexiang Liu, Zekun Li, Zhi Fang, Nan Xu 等ACL 2025 · 被引用 12 次
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order InformationRui Ai, Yuqi Pan, David Simchi-Levi, Milind Tambe 等ICML 2026 · 被引用 20 次
