Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
Feng Chen, Allan Raventós, Nan Cheng, Surya Ganguli, Shaul Druckmann
Abstract
Recent progress in large language models (LLMs) highlights the power of scaling test-time compute to achieve strong performance on complex tasks, such as mathematical reasoning and code generation. This raises a critical question: how should model training be modified to optimize performance under a subsequent test-time compute strategy and budget? To explore this, we focus on pass@N, a simple test-time strategy that searches for a correct answer in independent samples. We show, surprisingly, that training with cross-entropy (CE) loss can be with pass@N in that pass@N accuracy with longer training. We explain the origins of this misalignment in terms of model overconfidence induced by CE, and experimentally verify our prediction of overconfidence as an impediment to scaling test-time compute via pass@N. Furthermore we suggest a principled, modified training loss that is better aligned to pass@N by limiting model confidence and rescuing pass@N test performance. Our algorithm demonstrates improved mathematical reasoning on MATH and MiniF2F benchmarks under several scenarios: (1) providing answers to math questions; and (2) proving theorems by searching over proof trees of varying shapes. Overall our work underscores the importance of co-designing two traditionally separate phases of LLM development: training-time protocols and test-time search and reasoning strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7fea19d0-8de0-4eff-8d8c-a6cb537f30d5Cited by top-tier papers6
- Don’t Pass@k: A Bayesian Framework for Large Language Model EvaluationMohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin ChaudharyICLR 2026 · 18 citations
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time ScalingSachin Goyal, David Lopez-Paz, Kartik AhujaICLR 2026 · 11 citations
- Learning Shrinks the Hard Tail: Training‑Dependent Inference Scaling in a Solvable Linear ModelNoam Itzhak LeviICLR 2026 · 7 citations
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time ScalingIndranil Halder, Cengiz PehlevanICML 2026 · 4 citations
- Mode-conditioning unlocks superior test-time compute scalingChen Henry Wu, Sachin Goyal, Aditi RaghunathanICLR 2026 · 3 citations
Builds on23
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative VerifierJianyuan Zhong, Zeju Li, Zhijian Xu, Xiangyu Wen et al.ACL 2026 · 3 citations
- SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time ScalingYang Xiao, Chunpu Xu, Ruifeng Yuan, Jessie Wang et al.AAAI 2026 · 1 citation
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui et al.NeurIPS 2025 · 20 citations
- Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for ReasoningCharlie Victor Snell, Jaehoon Lee, Kelvin Xu, Aviral KumarICLR 2025
- EAGer: Entropy-Aware GEneRation for Adaptive Inference-Time ScalingDaniel Scalena, Leonidas Zotos, Elisabetta Fersini, Malvina Nissim et al.ICML 2026 · 13 citations
