On Predictability of Reinforcement Learning Dynamics for Large Language Models
Yuchen Cai, Ding Cao, Xin Xu, Zijun Yao, Yuqing Huang, Benyi Zhang, Zhenyu Tan, Guiquan Liu, Junfeng Fang
Abstract
Recent advances in reasoning capabilities of large language models (LLMs) are largely driven by reinforcement learning (RL), yet the underlying parameter dynamics during RL training remain poorly understood. This work identifies two fundamental properties of RL-induced parameter updates in LLMs: (1) Rank-1 Dominance, where the top singular subspace of the parameter update matrix nearly fully determines reasoning improvements, recovering over 99% of performance gains; and (2) Rank-1 Linear Dynamics, where this dominant subspace evolves linearly throughout training, enabling accurate prediction from early checkpoints. Extensive experiments across 8 LLMs and 7 algorithms validate the generalizability of these properties. More importantly, based on these findings, we propose AlphaRL, a plug-in acceleration framework that extrapolates the final parameter update using a short early training window, achieving up to 2.5× speedup while retaining >96% of reasoning performance without extra modules or hyperparameter tuning. This positions our finding as a versatile and practical tool for large-scale RL, opening a path toward principled, interpretable, and efficient training paradigm for LLMs. Our model and code will be available at: https://github.com/caiyuchenustc/Alpha- RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- The Emergence of Abstract Thought in Large Language Models Beyond Any LanguageYuxin Chen, Yiran Zhao, Yang Zhang, An Zhang et al.NeurIPS 2025 · 27 citations
- Reasoning Can Be Restored by Correcting a Few Decision TokensShen Changshuo, Leheng Sheng, Yuxin Chen, Xiang Wang et al.ICML 2026 · 2 citations
- GeoRA: Geometry-Aware Low-Rank Adaptation for RLVRJiaying Zhang, Lei Shi, Jiguo Li, Jun Xu et al.ACL 2026 · 1 citation
- FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision–Language ModelsChenyu Huang, Peng Ye, Xudong Tan, Jinhan Mu et al.ICML 2026
- Beyond Magnitude: Leveraging Direction of RLVR Updates for LLM ReasoningKexin Huang, Haoming Meng, Junkang Wu, Jinda Lu et al.ICLR 2026
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
Related papers
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu et al.NeurIPS 2025 · 273 citations
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng et al.ICLR 2026 · 4 citations
- Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern SelectionXingwu Chen, Tianle Li, Difan ZouICLR 2026 · 8 citations
- Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning ChainsWenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo et al.NeurIPS 2025 · 103 citations
- Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsSagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, Hao PengNeurIPS 2025 · 43 citations
