RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
Yiyou Sun, Yuhan Cao, Pohao Huang, Haoyue Bai, Hanna Hajishirzi, Nouha Dziri, Dawn Song
Abstract
It remains an open question whether LLMs can acquire or generalize genuinely new reasoning strategies, beyond the sharpened skills encoded in their parameters during pre-training or post-training. To attempt to answer this debate, we introduce DELTA -Distributional Evaluation of Learnability and Transferrability in Algorithmic Coding, a controlled benchmark of synthetic coding problem families designed to probe two fundamental aspects: learnability-can LLMs, through reinforcement learning (RL), solve problem families where pretrained models exhibit failure with large enough attempts (pass@K=0)?-and transferabilityif learnability happens, can such skills transfer systematically to out-of-distribution (OOD) test sets? Unlike prior public coding datasets, DELTA isolates reasoning skills through templated problem generators and introduces fully OOD problem families that demand novel strategies rather than tool invocation or memorized patterns. Our experiments reveal a striking grokking phase transition: after an extended period with near-zero reward, RL-trained models abruptly climb to near-perfect accuracy. To enable learnability on previously unsolvable problem families, we explore key training ingredients such as staged warm-up with dense rewards, experience replay, curriculum training, and verification-in-the-loop. Beyond learnability, we use DELTA to evaluate transferability or generalization along exploratory, compositional, and transformative axes, as well as cross-family transfer. Results show solid gains within families and for recomposed skills, but persistent weaknesses in transformative cases. DELTA thus offers a clean testbed for probing the limits of RL-driven reasoning and for understanding how models can move beyond existing priors to acquire new algorithmic skills. Code is available in https://github.com/sunblaze-ucb/rl-grok-recipe . Why controlled problem families matter? Uncontrolled open benchmarks in math/coding (e.g.,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c0c2a1f-aef8-4668-ab7b-443cf2b77e65Cited by top-tier papers7
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 58 citations
- Teaching Models to Teach Themselves: Reasoning at the Edge of LearnabilityShobhita Sundaram, John Quan, Ariel Kwiatkowski, Kartik Ahuja et al.ICML 2026 · 16 citations
- Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RLIan Wu, Yuxiao Qu, Amrith Setlur, Aviral KumarICML 2026 · 8 citations
- On the Emergence of Implicit Curriculum in RLVR Learning DynamicsYu Huang, Zixin Wen, Yuejie Chi, Yuting Wei et al.ICML 2026 · 6 citations
- The Unlearnability Phenomenon in RLVR for Language ModelsYulin Chen, He He, Chen ZhaoICML 2026 · 1 citation
Builds on6
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing ReasoningZhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu et al.ICLR 2026 · 271 citations
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsMingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu et al.NeurIPS 2025 · 181 citations
- Autoformalizing Euclidean GeometryLogan Murphy, Kaiyu Yang, Jialiang Sun, Zhaoyu Li et al.ICML 2024 · 16 citations
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsFangyu Lei, Jixuan Chen, Yuxiao Ye, Ruisheng Cao et al.ICLR 2025
- Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with TransformersRoman Abramov, Felix Steinbauer, Gjergji KasneciICML 2025
Related papers
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesLifan Yuan, Weize Chen, Yuchen Zhang, Ganqu Cui et al.ICLR 2026 · 46 citations
- SATURN: SAT-based Reinforcement Learning to Unleash LLMs ReasoningHuanyu Liu, Ge Li, Jia Li, Hao Zhu et al.NeurIPS 2025 · 1 citation
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningShubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling et al.ICLR 2026 · 112 citations
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement LearningYang Chen, Zhuolin Yang, Zihan Liu, Chankyu Lee et al.NeurIPS 2025 · 79 citations
