On Predictability of Reinforcement Learning Dynamics for Large Language Models
Yuchen Cai, Ding Cao, Xin Xu, Zijun Yao, Yuqing Huang, Benyi Zhang, Zhenyu Tan, Guiquan Liu, Junfeng Fang
摘要
Recent advances in reasoning capabilities of large language models (LLMs) are largely driven by reinforcement learning (RL), yet the underlying parameter dynamics during RL training remain poorly understood. This work identifies two fundamental properties of RL-induced parameter updates in LLMs: (1) Rank-1 Dominance, where the top singular subspace of the parameter update matrix nearly fully determines reasoning improvements, recovering over 99% of performance gains; and (2) Rank-1 Linear Dynamics, where this dominant subspace evolves linearly throughout training, enabling accurate prediction from early checkpoints. Extensive experiments across 8 LLMs and 7 algorithms validate the generalizability of these properties. More importantly, based on these findings, we propose AlphaRL, a plug-in acceleration framework that extrapolates the final parameter update using a short early training window, achieving up to 2.5× speedup while retaining >96% of reasoning performance without extra modules or hyperparameter tuning. This positions our finding as a versatile and practical tool for large-scale RL, opening a path toward principled, interpretable, and efficient training paradigm for LLMs. Our model and code will be available at: https://github.com/caiyuchenustc/Alpha- RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Emergence of Abstract Thought in Large Language Models Beyond Any LanguageYuxin Chen, Yiran Zhao, Yang Zhang, An Zhang 等NeurIPS 2025 · 被引用 27 次
- Reasoning Can Be Restored by Correcting a Few Decision TokensShen Changshuo, Leheng Sheng, Yuxin Chen, Xiang Wang 等ICML 2026 · 被引用 2 次
- GeoRA: Geometry-Aware Low-Rank Adaptation for RLVRJiaying Zhang, Lei Shi, Jiguo Li, Jun Xu 等ACL 2026 · 被引用 1 次
- FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision–Language ModelsChenyu Huang, Peng Ye, Xudong Tan, Jinhan Mu 等ICML 2026
- Beyond Magnitude: Leveraging Direction of RLVR Updates for LLM ReasoningKexin Huang, Haoming Meng, Junkang Wu, Jinda Lu 等ICLR 2026
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
相关 Paper
- AREAL: A Large-Scale Asynchronous Reinforcement Learning System for Language ReasoningWei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu 等NeurIPS 2025 · 被引用 273 次
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng 等ICLR 2026 · 被引用 4 次
- Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern SelectionXingwu Chen, Tianle Li, Difan ZouICLR 2026 · 被引用 8 次
- Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning ChainsWenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo 等NeurIPS 2025 · 被引用 103 次
- Reinforcement Learning Finetunes Small Subnetworks in Large Language ModelsSagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur, Hao PengNeurIPS 2025 · 被引用 43 次
