ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
Chu Zhao, Enneng Yang, Yuting Liu, Jianzhe Zhao, Guibing Guo
摘要
Test-time reinforcement learning generates multiple candidate answers via repeated rollouts and performs online updates using pseudo-labels constructed by majority voting. To reduce overhead and improve exploration, prior work introduces tree-structured rollouts, which share reasoning prefixes and branch at key nodes to improve sampling efficiency. However, this paradigm still faces two challenges: (1) high-entropy branching can trigger rollout collapse, where the branching budget concentrates on a few trajectories with consecutive high-entropy segments, rapidly reducing the number of effective branches; (2) early pseudo-labels are noisy and biased, which can induce self-reinforcing overfitting, causing the policy to sharpen prematurely and suppress exploration. To address these issues, we propose Entropy-Confidence Hybrid Group Relative Policy Optimization (ECHO). During rollout, ECHO jointly leverages local entropy and group-level confidence to adaptively control branch width, and further introduces online confidence-based pruning to terminate persistently low-confidence branches, avoiding high-entropy traps and mitigating collapse. During policy updates, ECHO employs confidence-adaptive clipping and an entropy-confidence hybrid advantage shaping approach to enhance training robustness and mitigate early-stage bias. Experiments demonstrate that ECHO achieves consistent gains on multiple mathematical and visual reasoning benchmarks, and generalizes more effectively under a limited rollout budget. The implementation code is available at https://github.com/ user683/ECHO
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang 等NeurIPS 2025 · 被引用 533 次
- Learning to Reason without External RewardsXuandong Zhao, Zhewei Kang, Aosong Feng, Sergey Levine 等ICLR 2026 · 被引用 218 次
- Deep Think with ConfidenceYichao Fu, Xuewei Wang, Hao Zhang, Yuandong Tian 等ICLR 2026 · 被引用 171 次
相关 Paper
- What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test TimeDong Yan, Jian Liang, Yanbo Wang, Shuo Lu 等ACL 2026 · 被引用 3 次
- Toward Generalized Web Agent Training: A Deep Dive into Entropy-Balanced Reinforcement LearningGuanting Dong, Licheng Bao, Zhongyuan Wang, Kangzhi Zhao 等WWW 2026 · 被引用 2 次
- From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image GenerationHan Song, Yucheng Zhou, Jianbing Shen, Yu ChengICLR 2026 · 被引用 9 次
- Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM ReasoningQiao Liang, Yuke Zhu, Chao Ge, Lei Yang 等ACL 2026 · 被引用 4 次
- Optimizing Anytime Reasoning via Budget Relative Policy OptimizationPenghui Qi, Zichen Liu, Tianyu Pang, Chao Du 等NeurIPS 2025 · 被引用 29 次
