One-Way Policy Optimization for Self-Evolving LLMs
Shuo Yang, Jinda Lu, Kexin Huang, Chiyu Ma, Shaohang Wei, Yuyang Liu, Guoyin Wang, Jingren Zhou, Li Yuan
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has become a promising paradigm for scaling reasoning capabilities of Large Language Models (LLMs). However, the sparsity of binary verifier rewards often leads to low efficiency and optimization instability. To stabilize training, existing methods typically impose token-level constraints relative to a reference policy. We identify that such constraints penalize deviations indiscriminately; this can flip verifier-determined direction when the policy attempts to outperform the reference, thereby suppressing gains. To resolve this, we propose One-Way Policy Optimization (OWPO) , a method based on the principle of decoupling optimization direction from update magnitude. In OWPO, the verifier dictates the update direction, while the reference policy serves only to adjust the magnitude. Specifically, OWPO applies asymmetric reweighting: it performs Accelerated Alignment for Inferior deviations (where the policy lags behind the reference) and Gain Locking for Superior deviations (where the policy surpasses the reference). Furthermore, by incorporating iterative reference updates, OWPO creates a ``Ratchet Effect'' that continuously consolidates gains. Experimental results demonstrate that OWPO outperforms strong baselines, including DAPO, OPD, and MOPD, breaking the bottleneck of fixed priors to enable continuous self-evolution without reliance on external reference models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4ff1f29-8744-43ca-8ca1-cb8d83f11158Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Beyond Magnitude: Leveraging Direction of RLVR Updates for LLM ReasoningKexin Huang, Haoming Meng, Junkang Wu, Jinda Lu et al.ICLR 2026
- Rethinking Sample Polarity in Reinforcement Learning with Verifiable RewardsXinyu Tang, Yuliang Zhan, Zhixun Li, Xin Zhao et al.ACL 2026 · 20 citations
- Well Begun, Half Done: Reinforcement Learning with Prefix Optimization for LLM ReasoningYiliu Sun, Zicheng Zhao, Yang Wei, Yanfang Zhang et al.AAAI 2026 · 1 citation
- TGPO: Efficient Policy Optimization through Sequence Anchor and Information GatingHang Ding, Dongqi Liu, Qiming Feng, Jian Li et al.ICML 2026
- MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM ReasoningXiaoliang Fu, Jiaye Lin, Yangyi Fang, Binbin Zheng et al.ACL 2026 · 12 citations
