Self-Improvement in Language Models: The Sharpening Mechanism
Audrey Huang, Adam Block, Dylan J. Foster, Dhruv Rohatgi, Cyril Zhang, Max Simchowitz, Jordan T. Ash, Akshay Krishnamurthy
Abstract
Recent work in language modeling has raised the possibility of self-improvement, where a language models evaluates and refines its own generations to achieve higher performance without external feedback. It is impossible for this self-improvement to create information that is not already in the model, so why should we expect that this will lead to improved capabilities? We offer a new perspective on the capabilities of self-improvement through a lens we refer to as sharpening. Motivated by the observation that language models are often better at verifying response quality than they are at generating correct responses, we formalize self-improvement as using the model itself as a verifier during post-training in order to "sharpen" the model to one placing large mass on high-quality sequences, thereby amortizing the expensive inference-time computation of generating good sequences. We begin by introducing a new statistical framework for sharpening in which the learner aims to sharpen a pre-trained base policy via sample access, and establish fundamental limits. Then, we analyze two natural families of self-improvement algorithms based on SFT and RLHF. We find that (i) the SFT-based approach is minimax optimal whenever the initial model has sufficient coverage, but (ii) the RLHF-based approach can improve over SFT-based self-improvement by leveraging online exploration, bypassing the need for coverage. Finally, we empirically validate the sharpening mechanism via inference-time and amortization experiments. We view these findings as a starting point toward a foundational understanding that can guide the design and evaluation of self-improvement algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a746d00-243e-4bd1-977b-43bfc5ae0b25Cited by top-tier papers46
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine et al.ICLR 2026 · 218 citations
- The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningShivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han et al.NeurIPS 2025 · 185 citations
- Self-Adapting Language ModelsAdam Zweiger, Jyothish Pari, Han Guo, Yoon Kim et al.NeurIPS 2025 · 78 citations
- Reinforcing General Reasoning Without VerifiersXiangxin Zhou, Zichen Liu, Anya Sims, Haonan Wang et al.ICLR 2026 · 75 citations
- e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMsAmrith Setlur, Matthew Y. R. Yang, Charlie Victor Snell, Jeremiah Greer et al.ICLR 2026 · 66 citations
Builds on45
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language ModelsYuda Song, Hanlin Zhang, Carson Eisenach, Sham M. Kakade et al.ICLR 2025 · 3 citations
- Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution SharpeningXiaotong Ji, Rasul Tutunov, Matthieu Zimmer, Haitham Bou AmmarICML 2026 · 17 citations
- Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time AlignmentAudrey Huang, Adam Block, Qinghua Liu, Nan Jiang et al.ICML 2025
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui et al.NeurIPS 2025 · 20 citations
- Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier GapYifan Sun, Yushan Liang, Zhen Zhang, Xin Liu et al.ICLR 2026 · 2 citations
