Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
Amirhosein Ghasemabadi, Keith G. Mills, Baochun Li, Di Niu
Abstract
Test-Time Scaling (TTS) methods for enhancing Large Language Model (LLM) reasoning often incur substantial computational costs, primarily due to extensive reliance on external Process Reward Models (PRMs) or sampling methods like Best-of-N (BoN). This paper introduces Guided by Gut (GG), an efficient self-guided TTS framework that achieves PRM-level performance without costly external verifier models. Our method employs a lightweight tree search guided solely by intrinsic LLM signals, token-level confidence and step novelty. One critical innovation is improving the reliability of internal confidence estimates via a targeted reinforcement learning fine-tuning phase. Empirical evaluations on challenging mathematical reasoning benchmarks demonstrate that GG enables smaller models (e.g., 1.5B parameters) to achieve accuracy matching or surpassing significantly larger models (e.g., 32B-70B parameters), while reducing GPU memory usage by up to 10x. Compared to PRM-based methods, GG achieves comparable accuracy with 8x faster inference speeds and 4-5x lower memory usage. Additionally, GG reduces KV cache memory usage by approximately 50% compared to the BoN strategy, facilitating more efficient and practical deployment of TTS techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PRInTS: Reward Modeling for Long-Horizon Information SeekingJaewoo Lee, Archiki Prasad, Justin Chih-Yao Chen, Zaid Khan et al.ACL 2026 · 3 citations
- CaTS: Calibrated Test-Time Scaling for Efficient LLM ReasoningChengsong Huang, Langlin Huang, Jixuan Leng, Jiacheng Liu et al.ICLR 2026
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language ModelsJingwei Ni, Ekaterina Fadeeva, Tianyi Wu, Mubashara Akhtar et al.ACL 2026 · 1 citation
- Step-level Verifier-guided Hybrid Test-Time Scaling for Large Language ModelsKaiyan Chang, Yonghao Shi, Chenglong Wang, Hang Zhou et al.EMNLP 2025
- Confidence-Guided Stepwise Model Routing for Cost-Efficient ReasoningSangmook Lee, Dohyung Kim, Hyukhun Koh, Nakyeong Yang et al.AAAI 2026 · 3 citations
- MUR: Momentum Uncertainty guided Reasoning for Large Language ModelsHang Yan, Fangzhi Xu, Rongman Xu, Yifei Li et al.ACL 2026 · 12 citations
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui et al.NeurIPS 2025 · 20 citations
