ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
Ruiyang Zhou, Shuozhe Li, Amy Zhang, Liu Leqi
摘要
Self-improvement via RL often fails on complex reasoning tasks because GRPO-style post-training methods rely on the model's initial ability to generate positive samples. Without guided exploration, these approaches merely reinforce what the model already knows (distribution-sharpening) rather than enabling the model to solve problems where it initially generates no correct solutions. To unlock reasoning ability in such settings, the model must explore new reasoning trajectories beyond its current output distribution. Such exploration requires access to sufficiently good positive samples to guide the learning. While expert demonstrations seem like a natural solution, we find that they are often ineffective in RL post-training. Instead, we identify two key properties of effective positive samples: they should (1) be likely under the current policy, and (2) increase the model's likelihood of predicting the correct answer. Based on these insights, we propose -a simple and modular framework that generates such samples by conditioning on the ground-truth answer. It can be integrated with popular RL training methods like GRPO and DPO. ExPO enables efficient exploration and guides the model to produce reasoning trajectories more aligned with its policy than expert-written CoTs, while ensuring higher quality than its own (incorrect) samples. Experiments show that ExPO improves both learning efficiency and final performance on reasoning benchmarks, surpassing expert-demonstration-based methods in challenging settings such as MATH level-5, where the model initially struggles the most. Code is available at https://github.com/HumainLab/ExPO_rl_reasoning_by_explanation .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Nudging the Boundaries of LLM ReasoningJustin Chih-Yao Chen, Xiangyu Peng, Prafulla Kumar Choubey, Kung-Hsiang Huang 等ICLR 2026 · 被引用 25 次
- Representation-Based Exploration for Language Models: From Test-Time to Post-TrainingJens Tuyls, Dylan J Foster, Akshay Krishnamurthy, Jordan T. AshICLR 2026 · 被引用 18 次
- Off-Trajectory Reasoning: Can LLMs Collaborate on Reasoning Trajectories?Aochong Oliver Li, Tanya GoyalICLR 2026 · 被引用 6 次
- Chart Deep Research in LVLMs via Parallel Relative Policy OptimizationJiajin Tang, Yang Gao, Wenjie Wang, Sibei Yang 等ICLR 2026 · 被引用 1 次
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann 等ICML 2026
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree SearchDan Zhang, Sining Zhoubian, Ziniu Hu, Yisong Yue 等NeurIPS 2024 · 被引用 527 次
相关 Paper
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng 等ICML 2026 · 被引用 17 次
- Trust-Region Adaptive Policy OptimizationMingyu Su, Jian Guan, Yuxian Gu, Minlie Huang 等ICLR 2026 · 被引用 2 次
- Better, Faster: Harnessing Self-Improvement in Large Reasoning ModelsQihuang Zhong, Liang Ding, Juhua Liu, Bo Du 等ICML 2026 · 被引用 3 次
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng 等ICLR 2026 · 被引用 4 次
- RESTRAIN: From Spurious Votes to Signals - Self-Training RL with Self-PenalizationZhaoning Yu, Zhaolun Su, Leitian Tao, Haozhu Wang 等ICLR 2026 · 被引用 7 次
