Experience Augmented Policy Optimization for LLM Reasoning
Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou
摘要
Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing RLVR methods typically rely on on-policy optimization from scratch, resulting in high sampling costs and inefficient utilization of accumulated experience. As model capabilities and policy behaviors evolve during training, recent attempts to reuse experience via fixed reasoning trajectories further suffer from policy mismatch. Motivated by these limitations, we argue that experience in RLVR should not be reused as fixed reasoning trajectories, but instead expressed in a policy-adaptive manner. In this work, we propose Experience-Augmented Policy Optimization (EAPO), which leverages a prior RL-optimized policy as an action-level experience prior and selectively injects experience at critical decision points during rollout. To ensure stable and unbiased learning from experience-augmented rollouts, EAPO further incorporates an adapted importance sampling scheme. Experiments on using Qwen-2.5-math 7b and Qwen-3-8B on five different benchmarks demonstrate that EAPO consistently improves reasoning performance over state-of-the-art RLVR methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin 等AAAI 2024 · 被引用 484 次
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
相关 Paper
- ExGRPO: Learning to Reason from ExperienceRunzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu 等ICLR 2026 · 被引用 51 次
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-ForgetLiang Chen, Xueting Han, Qizhou Wang, Bo Han 等ICLR 2026 · 被引用 16 次
- FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable ReasoningYuyang Ding, Chi Zhang, Juntao Li, Haibin Lin 等ICLR 2026 · 被引用 7 次
- EAPO: Enhancing Policy Optimization with On-Demand Expert AssistanceSiyao Song, Cong Ma, Zhihao Cheng, Shiye Lei 等ICML 2026
- From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive OptimizationXinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li 等ACL 2026 · 被引用 9 次
