TTRL: Test-Time Reinforcement Learning
Yuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu, Ganqu Cui, Xuekai Zhu, Haozhan Li, Yuchen Zhang, Xinwei Long, Ermo Hua, Biqing Qi, Youbang Sun
Abstract
This paper investigates Reinforcement Learning (RL) on data without explicit labels for reasoning tasks in Large Language Models (LLMs). The core challenge of the problem is reward estimation during inference while not having access to ground-truth information. While this setting appears elusive, we find that common practices in Test-Time Scaling (TTS), such as majority voting, yield surprisingly effective rewards suitable for driving RL training. In this work, we introduce Test-Time Reinforcement Learning (TTRL), a novel method for training LLMs using RL on unlabeled data. TTRL enables self-evolution of LLMs by utilizing the priors in the pre-trained models. Our experiments demonstrate that TTRL consistently improves performance across a variety of tasks and models. Notably, TTRL boosts the pass@1 performance of Qwen-2.5-Math-7B by approximately 211% on the AIME 2024 with only unlabeled test data. Furthermore, although TTRL is only supervised by the maj@n metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model maj@n, and approach the performance of models trained directly on test data with ground-truth labels. Our experimental findings validate the general effectiveness of TTRL across various tasks and highlight TTRL's potential for broader tasks and domains. GitHub: https://github.com/PRIME-RL/TTRL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c1a0a05-bc46-429b-8dff-5e8a81ab135cCited by top-tier papers54
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang et al.NeurIPS 2025 · 310 citations
- R-Zero: Self-Evolving Reasoning LLM from Zero DataChengsong Huang, Wenhao Yu, Xiaoyang Wang, Hongming Zhang et al.ICLR 2026 · 220 citations
- The Unreasonable Effectiveness of Entropy Minimization in LLM ReasoningShivam Agarwal, Zimin Zhang, Lifan Yuan, Jiawei Han et al.NeurIPS 2025 · 185 citations
- Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning IncentivizationQingyang Zhang, Haitao Wu, Changqing Zhang, Peilin Zhao et al.NeurIPS 2025 · 134 citations
Builds on30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Reinforcement Learning on Pre-Training DataSiheng Li, Kejiao Li, Zenan Xu, Guanhua Huang et al.ACL 2026 · 11 citations
- Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM ReasoningSitong Wu, Haoru Tan, Xichen Zhang, Bin Xia et al.ICML 2026
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong et al.ICLR 2026 · 42 citations
- Incentivizing LLMs to Self-Verify Their AnswersFuxiang Zhang, Jiacheng Xu, Chaojie Wang, Ce Cui et al.NeurIPS 2025 · 20 citations
- What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test TimeDong Yan, Jian Liang, Yanbo Wang, Shuo Lu et al.ACL 2026 · 3 citations
