Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
Jingkai Huang, Will Ma, Zhengyuan Zhou
摘要
A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached. In this paper we leverage Bayesian prior information to save on sampling costs, stopping once sufficient consistency is reached. Although the exact posterior is computationally intractable, we further introduce an efficient ``-aggregated'' stopping policy that tracks only the most frequent answer counts. Theoretically, we prove that is all you need: this coarse approximation is sufficient to achieve asymptotic optimality, and strictly dominates prior-free baselines, while having a fast posterior computation. Empirically, this identifies the most consistent (i.e., mode) LLM answer and achieves similar answer accuracy using fewer samples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Are More LLM Calls All You Need? Towards the Scaling Properties of Compound AI SystemsLingjiao Chen, Jared Quincy Davis, Boris Hanin, Peter Bailis 等NeurIPS 2024 · 被引用 110 次
- Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step ReasoningYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan 等ICLR 2024 · 被引用 101 次
相关 Paper
- Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsPranjal Aggarwal, Aman Madaan, Yiming Yang, MausamEMNLP 2023 · 被引用 5 次
- Optimal Self-Consistency for Efficient Reasoning with Large Language ModelsAustin Feng, Marius Alonso, Ambroise Odonnat, Vasilii Feofanov 等ICML 2026 · 被引用 6 次
- Statistical Early Stopping for Reasoning ModelsYangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun 等ICML 2026 · 被引用 4 次
- Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time ComputeJianhao Chen, Zishuo Xun, Bocheng Zhou, Han Qi 等AAAI 2026 · 被引用 18 次
- From Drift to Coherence: Stabilizing Beliefs in LLMsSongEun Kim, Seungyoo Lee, Edwin Fong, Hyungi Lee 等ICML 2026 · 被引用 2 次
