GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
Divij Handa, Mihir Parmar, Aswin RRV, Md Nayem Uddin, Hamid Palangi, Chitta Baral
摘要
Repeated Sampling (RS) is a simple inference-time algorithm that has been shown to improve model performance on complex tasks. Although it is an effective way of scaling inference time, it often struggles to generate diverse solution candidates, frequently relying on the same underlying approach to solve the problem and thus producing redundant samples. To address this limitation, we propose a new inference algorithm, GUIDEDSAMPLING, which decouples the exploration and generation phases during inference, increasing diversity of generated candidate solutions. The exploration phase identifies multiple concepts that can be utilized to solve the problem, while the generation phase applies a specific concept to provide final solution candidates. We first define the theoretical bounds of GUID-EDSAMPLING and then empirically demonstrate that it improves the performance of base model at pass@50 by on an average ∼ 21.6% across various benchmarks compared to RS. Furthermore, models trained on trajectories of GUIDEDSAM-PLING exhibit substantial performance improvements at pass@5 by on an average ∼ 9.7%, compared to models trained on traditional RS. Additionally, models trained with GUIDEDSAMPLING increases the average number of concepts per instance (1.67 → 3.03), yielding a diverse set of candidates than traditional RS. 1 Recently, various inference-time algorithms have been proposed (
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Scaling Group Inference for Diverse and High-Quality GenerationGaurav Parmar, Or Patashnik, Daniil Ostashev, Kuan-Chieh Wang 等ICLR 2026 · 被引用 14 次
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree SearchYuichi Inoue, Kou Misaki, Yuki Imajuku, So Kuroki 等NeurIPS 2025 · 被引用 67 次
- Better, Faster: Harnessing Self-Improvement in Large Reasoning ModelsQihuang Zhong, Liang Ding, Juhua Liu, Bo Du 等ICML 2026 · 被引用 3 次
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng 等ICML 2026 · 被引用 21 次
- Planning in Natural Language Improves LLM Search for Code GenerationEvan Z. Wang, Federico Cassano, Catherine Wu, Yunfeng Bai 等ICLR 2025
