h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement Learning
Alesia Ivanova, Sumeet Motwani, Jack Cai, Phil Torr, Riashat Islam, Shital Shah, Christian Schroeder de Witt, Charles London
Abstract
Large language models excel at short-horizon reasoning tasks, but performance drops as reasoning horizon lengths increase. Existing approaches to combat this rely on inference-time scaffolding or costly step-level supervision, neither of which scales easily. In this work, we introduce a scalable method to bootstrap long-horizon reasoning capabilities using only existing, abundant short-horizon data. Our approach synthetically composes simple problems into complex, multistep dependency chains of arbitrary length. We train models on this data using outcome-only rewards under a curriculum that automatically increases in complexity, allowing RL training to be scaled much further without saturating. Empirically, our method generalizes remarkably well: curriculum training on composed 6th-grade level math problems (GSM8K) boosts accuracy on longer, competitionlevel benchmarks (GSM-Symbolic, MATH-500, AIME) by up to 2.06×. It also transfers significantly to diverse out-of-distribution ReasoningGym domains and long-context benchmarks, indicating broader generalization. Importantly, our long-horizon improvements are significantly higher than baselines even at high pass@k, showing that models can learn new reasoning paths under RL. Theoretically, we show that curriculum RL with outcome rewards achieves an exponential improvement in sample complexity over full-horizon training, providing training signal comparable to dense supervision. h1 therefore introduces an efficient path towards scaling RL for long-horizon problems using only existing data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab1ad738-7cd0-4785-b399-810d5a4c0eeeCited by top-tier papers2
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song et al.ICML 2026 · 18 citations
- LongCoT: Benchmarking Long-Horizon Chain-of-Thought ReasoningSumeet Motwani, Daniel Nichols, Charles London, Peggy Li et al.ICML 2026 · 2 citations
Builds on16
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton et al.NeurIPS 2025 · 507 citations
- Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleYiping Wang, Qing Yang, Zhiyuan Zeng, Liliang Ren et al.NeurIPS 2025 · 314 citations
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
Related papers
- RL for Reasoning by Adaptively Revealing RationalesMohammad Hossein Amani, Aryo Lotfi, Nicolas Baldwin, Samy Bengio et al.ICLR 2026 · 19 citations
- Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement LearningZhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin et al.ICML 2024 · 70 citations
- R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?Yi Lu, Jianing Wang, Linsen Guo, Wei He et al.ICLR 2026 · 9 citations
- GSM-∞: How Do your LLMs Behave over Infinitely Increasing Reasoning Complexity and Context Length?Yang Zhou, Hongyi Liu, Zhuoming Chen, Yuandong Tian et al.ICML 2025
- Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement LearningWeijie Li, Jin Wang, Liang-Chih Yu, Xuejie ZhangAAAI 2026 · 1 citation
