Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning
Weiqin Wang, Yile Wang, Kehao Chen, Hui Huang
摘要
Test-time reinforcement learning mitigates the reliance on annotated data by using majority voting results as pseudo-labels, emerging as a complementary direction to reinforcement learning with verifiable rewards (RLVR) for improving reasoning ability of large language models (LLMs). However, this voting strategy often induces confirmation bias and suffers from sparse rewards, limiting the overall performance. In this work, we propose subgroup-specific step-wise confidenceweighted pseudo-label estimation (SCOPE), a framework integrating model confidence and dynamic subgroup partitioning to address these issues. Specifically, SCOPE integrates the proposed step-wise confidence into pseudolabel estimation, prioritizing high-quality reasoning paths over simple frequency count. Furthermore, it dynamically partitions the candidate outputs pool into independent subgroups by balancing reasoning quality against exploration diversity. By deriving local consensus via repeat sampling for each subgroup, SCOPE provides diverse supervision targets to encourage broader exploration. We conduct experiments across various models and benchmarks, experimental results show that SCOPE consistently outperforms recent baselines. Notably, SCOPE achieves relative improvements of 13.1% on challenging AIME 2025 and 8.1% on AMC. The code is released at https://github.com/szu-tera/SCOPE . r(o 1 ,o * ) r(o 2 ,o * ) r(o 3 ,o * ) r(o 4 ,o * ) r(o 5 ,o * ) r(o 6 ,o * ) r(o 7 ,o * ) r(o 8 ,o * ) subgroup partition o 1 , o 2 o 3 , o 4 o 5 , o 6 o 7 , o 8 subgroup size = 2 different consensus for each subgroup r(o 1 ,o 1 * ) r(o 3 ,o 2 * ) r(o 5 ,o 3 * ) r(o 7 ,o 4 * ) r(o 2 ,o 1 * ) r(o 4 ,o 2 * ) r(o 6 ,o 3 * ) r(o 8 ,o 4 * )
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal StructureZirui Li, Xuefeng Bai, Kehai Chen, Yizhi Li 等ICML 2026
- Diagnosing and Remedying Representation Deficiencies for Deterministic Reasoning in KGQAGewen Liang, Mufan Xu, Kehai Chen, Wei Wang 等ACL 2026
它引用的顶会 Paper14
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye 等ICLR 2026 · 被引用 279 次
- TTRL: Test-Time Reinforcement LearningYuxin Zuo, Kaiyan Zhang, Li Sheng, Shang Qu 等NeurIPS 2025 · 被引用 249 次
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine 等ICLR 2026 · 被引用 218 次
相关 Paper
- Beyond Majority Voting: Self-Reflective Test-Time Reinforcement Learning for LLM ReasoningSitong Wu, Haoru Tan, Xichen Zhang, Bin Xia 等ICML 2026
- What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test TimeDong Yan, Jian Liang, Yanbo Wang, Shuo Lu 等ACL 2026 · 被引用 3 次
- ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement LearningChu Zhao, Enneng Yang, Yuting Liu, Jianzhe Zhao 等ICML 2026 · 被引用 2 次
- The Surprising Effectiveness of Negative Reinforcement in LLM ReasoningXinyu Zhu, Mengzhou Xia, Zhepei Wei, Wei-Lin Chen 等NeurIPS 2025 · 被引用 177 次
- d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language ModelsLeyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu 等ACL 2026 · 被引用 14 次
