Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM Reasoning
Sitong Wu, Haoru Tan, Xichen Zhang, Bin Xia, Shaofeng Zhang, XIAOJUAN QI, Bei Yu, Jiaya Jia
Abstract
The core challenge of Test-Time Reinforcement Learning (TTRL) lies in estimating rewards without access to ground-truth supervision. Existing TTRL methods predominantly rely on majority voting to generate pseudo-labels, under the assumption that the most frequent answer among sampled trajectories is correct. However, we observe that this assumption frequently breaks down in complex reasoning tasks, where correct solutions often constitute a logical minority. As a result, rare yet correct trajectories are systematically undervalued by majority-voting-based approaches. To address this limitation, we propose Self-Reflective Test-Time Reinforcement Learning (SR-TTRL), a novel framework that leverages self-reflective verification to produce high-fidelity pseudo-labels. Specifically, given multiple sampled trajectories for a problem, SR-TTRL first groups trajectories according to their final answers and selects one representative from each group to form a candidate pool. Each candidate trajectory is then summarized to preserve its core reasoning steps while reducing verbosity. Finally, the model performs self-reflection over the candidate pool, critically evaluating and selecting the most plausible trajectory as the pseudo-label. Empirically, SR-TTRL achieves substantially higher pseudo-label fidelity and sample efficiency than prior majority-voting-based TTRL methods. Extensive experiments across diverse benchmarks and model families demonstrate that SR-TTRL consistently outperforms majority-voting baselines and significantly improves generalization to novel problems. For example, SR-TTRL improves the Pass@1 accuracy of Qwen3-8B on AIME24 from 29.1 to 55.8 (a gain of +26.7), exceeding standard TTRL by an additional +9.1. The code will be released at: https://github.com/JIA-Lab-research/SR-TTRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Test-Time Training with Self-Supervision for Generalization under Distribution ShiftsYu Sun, Xiaolong Wang, Zhuang Liu, John Miller et al.ICML 2020 · 1,220 citations
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
Related papers
- What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test TimeDong Yan, Jian Liang, Yanbo Wang, Shuo Lu et al.ACL 2026 · 3 citations
- Self-Harmony: Learning to Harmonize Self-Supervision and Self-Play in Test-Time Reinforcement LearningRu Wang, Wei Huang, Qi Cao, Yusuke Iwasawa et al.ICLR 2026 · 9 citations
- Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement LearningWeiqin Wang, Yile Wang, Kehao Chen, Hui HuangACL 2026 · 5 citations
- CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-EvolutionTeng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang et al.ACL 2026 · 5 citations
- S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement LearningRuotian Ma, Peisong Wang, Cheng Liu, Xingyan Liu et al.ACL 2025 · 13 citations
