Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning
Weiqin Wang, Yile Wang, Kehao Chen, Hui Huang
Abstract
Test-time reinforcement learning mitigates the reliance on annotated data by using majority voting results as pseudo-labels, emerging as a complementary direction to reinforcement learning with verifiable rewards (RLVR) for improving reasoning ability of large language models (LLMs). However, this voting strategy often induces confirmation bias and suffers from sparse rewards, limiting the overall performance. In this work, we propose subgroup-specific step-wise confidenceweighted pseudo-label estimation (SCOPE), a framework integrating model confidence and dynamic subgroup partitioning to address these issues. Specifically, SCOPE integrates the proposed step-wise confidence into pseudolabel estimation, prioritizing high-quality reasoning paths over simple frequency count. Furthermore, it dynamically partitions the candidate outputs pool into independent subgroups by balancing reasoning quality against exploration diversity. By deriving local consensus via repeat sampling for each subgroup, SCOPE provides diverse supervision targets to encourage broader exploration. We conduct experiments across various models and benchmarks, experimental results show that SCOPE consistently outperforms recent baselines. Notably, SCOPE achieves relative improvements of 13.1% on challenging AIME 2025 and 8.1% on AMC. The code is released at https://github.com/szu-tera/SCOPE . r(o 1 ,o * ) r(o 2 ,o * ) r(o 3 ,o * ) r(o 4 ,o * ) r(o 5 ,o * ) r(o 6 ,o * ) r(o 7 ,o * ) r(o 8 ,o * ) subgroup partition o 1 , o 2 o 3 , o 4 o 5 , o 6 o 7 , o 8 subgroup size = 2 different consensus for each subgroup r(o 1 ,o 1 * ) r(o 3 ,o 2 * ) r(o 5 ,o 3 * ) r(o 7 ,o 4 * ) r(o 2 ,o 1 * ) r(o 4 ,o 2 * ) r(o 6 ,o 3 * ) r(o 8 ,o 4 * )
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6fc03cd-18a6-4fde-b10b-43ad8a8ef76eCited by top-tier papers2
- Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal StructureZirui Li, Xuefeng Bai, Kehai Chen, Yizhi Li et al.ICML 2026
- Diagnosing and Remedying Representation Deficiencies for Deterministic Reasoning in KGQAGewen Liang, Mufan Xu, Kehai Chen, Wei Wang et al.ACL 2026
Builds on14
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye et al.ICLR 2026 · 279 citations
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu et al.NeurIPS 2025 · 249 citations
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine et al.ICLR 2026 · 218 citations
Related papers
- Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM ReasoningSitong Wu, Haoru Tan, Xichen Zhang, Bin Xia et al.ICML 2026
- What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test TimeDong Yan, Jian Liang, Yanbo Wang, Shuo Lu et al.ACL 2026 · 3 citations
- ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement LearningChu Zhao, Enneng Yang, Yuting Liu, Jianzhe Zhao et al.ICML 2026 · 2 citations
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen et al.NeurIPS 2025 · 177 citations
- d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language ModelsLeyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu et al.ACL 2026 · 14 citations
