Do Not Step Into the Same River Twice: Learning to Reason from Trial and Error
Chenming Tang, Hsiu-Yuan Huang, Weijie Liu, Clive Bai, Saiyong Yang, Yunfang Wu
摘要
Reinforcement learning with verifiable rewards (RLVR) has significantly boosted the reasoning capability of language models (LMs). However, existing RLVR approaches train LMs based on their own on-policy responses and are constrained by the initial capability of LMs, thus prone to exploration stagnation, in which LMs fail to solve more training problems and cannot further learn from the training data. Some approaches try to address this by leveraging off-policy solutions to training problems, but rely on external expert guidance that is limited in availability and scalability. In this work, we propose LTE (Learning to reason from Trial and Error), an approach that hints LMs with their previously self-made mistakes, not requiring any external expert guidance. Experiments validate the effectiveness of LTE, which outperforms the normal group relative policy optimization (GRPO) by 5.02 in Pass@1 and 9.96 in Pass@k on average across six mathematical reasoning benchmarks for Qwen3-8B-Base and even performs better than methods that require external guidance. Further analysis confirms that LTE successfully mitigates exploration stagnation and enhances both exploitation and exploration during training. Our code is available at https://github.com/ JamyDon/LTE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- No More Stale Feedback: Co-Evolving Critics for Open-World Agent LearningZhicong Li, Lingjie Jiang, Yulan Hu, Xingchen Zeng 等ACL 2026 · 被引用 3 次
- CURE: Critique-Driven Unified Reinforcement Learning for Test-Time Self-ImprovementGuirong Chen, Shuqi Ye, Wenkai Yang, Shiqi Shen 等ACL 2026
它引用的顶会 Paper11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic WeightingWenhao Zhang, Yuexiang Xie, Yuchang Sun, Yanxi Chen 等ICLR 2026 · 被引用 100 次
- Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data ContaminationMingqi Wu, Zhihao Zhang, Qiaole Dong, Zhiheng Xi 等AAAI 2026 · 被引用 67 次
相关 Paper
- StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to ReasonKaiyi Zhang, Ang Lv, Jinpeng Li, Yongbo Wang 等ACL 2026 · 被引用 32 次
- Experience Augmented Policy Optimization for LLM ReasoningJinda Lu, Kexin Huang, Junkang Wu, Shuo Yang 等ICML 2026 · 被引用 2 次
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM ReasoningXichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan 等ICLR 2026 · 被引用 52 次
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu 等ICLR 2026 · 被引用 51 次
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy OptimizationYihong Dong, Xue Jiang, Yongding Tao, Huanyu Liu 等ACL 2026 · 被引用 34 次
