HyPER: Bridging Exploration and Exploitation for Scalable LLM Reasoning with Hypothesis Path Expansion and Reduction
Shengxuan Qiu, Haochen Huang, Shuzhang Zhong, Pengfei Zuo, Meng Li
Abstract
Scaling test-time compute with multi-path chainof-thought can improve reasoning accuracy, but its gains hinge on an effective explorationexploitation trade-off. Existing methods handle this trade-off in rigid ways: tree-structured search hard-codes exploration via brittle expansion rules that disrupt post-trained reasoning, while parallel reasoning over-explores redundant hypothesis paths and relies on a weak answer selection strategy. Driven by the insight that the optimal balance is phase-dependent and that correct vs. incorrect paths often diverge only at late stages, we reconceptualize test-time scaling as a dynamic expand-reduce control problem over a pool of hypothesis paths. We introduce Hy-PER, a training-free online control policy for MoE multi-path decoding that reallocates compute under a fixed budget using lightweight path statistics. HyPER features (i) an online controller that shifts from exploration to exploitation as the hypothesis pool evolves, (ii) an MoE-based tokenlevel refinement primitive for efficient generationtime exploitation without full-path resampling, and (iii) a length-and confidence-aware aggregation rule to bridge the existence-selection gap for reliable answer-time exploitation. Extensive experimental results across four MoE models and diverse benchmarks demonstrate HyPER consistently achieves the accuracy-compute Pareto frontier, outperforming prior-art methods by 8 ∼ 10% while reducing token consumption by 25 ∼ 40%. Code is available at https://github.com/ ShengxuanQiu/HyPER .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f67680c-b253-4335-8ae9-423deac3cbfdBuilds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger et al.AAAI 2024 · 1,292 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- What If We Allocate Test-Time Compute Adaptively?Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer, Ali Subhan et al.ICML 2026 · 3 citations
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMsAmrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer et al.ICLR 2026 · 66 citations
- Parallel-Probe: Towards Efficient Parallel Thinking via 2D ProbingTong Zheng, Chengsong Huang, Runpeng Dai, Yun He et al.ICML 2026
- Optimizing Test-Time Compute via Meta Reinforcement FinetuningYuxiao Qu, Matthew Y. R. Yang, Amrith Setlur, Lewis Tunstall et al.ICML 2025
- Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought ReasoningRenos Zabounidis, Aditya Golatkar, Michael Kleinman, Alessandro Achille et al.ICML 2026 · 4 citations
