R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
Yi Lu, Jianing Wang, Linsen Guo, Wei He, Hongyin Tang, Tao Gui, Xuanjing Huang, Xuezhi Cao, Wei Wang, Xunliang Cai
摘要
Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek-R1) have led to remarkable improvements through long Chain-of-Thought (CoT). However, existing benchmarks mainly focus on immediate, single-horizon tasks, failing to adequately evaluate models’ ability to understand and respond to complex, long-horizon scenarios. To address this incomplete evaluation of Large Reasoning Models (LRMs), we propose R-HORIZON, a method designed to stimulate long-horizon reasoning behaviors in LRMs through query composition. Based on R-HORIZON, we construct a long-horizon reasoning benchmark, comprising complex multi-step reasoning tasks with interdependent problems that span long reasoning horizons. Through comprehensive evaluation of LRMs using the R-HORIZON benchmark, we find that even the most advanced LRMs suffer significant performance degradation. Our analysis reveals that LRMs exhibit limited effective reasoning length and struggle to allocate thinking budget across multiple problems appropriately. Recognizing these limitations, we use R-HORIZON to construct long-horizon reasoning data for reinforcement learning with verified rewards (RLVR). Compared to training with single-horizon data, RLVR with R-HORIZON not only substantially improves performance on the multi-horizon reasoning tasks, but also promotes accuracy on standard reasoning tasks (+7.5 on AIME2024). These results position R-HORIZON as a scalable, controllable, and low-cost paradigm for enhancing and evaluating the long-horizon reasoning capabilities of LRMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing ReasoningZhiwei He, Tian Liang, Jiahao Xu, Qiuzhi Liu 等ICLR 2026 · 被引用 271 次
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 被引用 270 次
- OpenThoughts: Data Recipes for Reasoning ModelsEtash Kumar Guha, Ryan Marten, Sedrick Keh, Negin Raoof 等ICLR 2026 · 被引用 235 次
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du 等ICLR 2026 · 被引用 225 次
相关 Paper
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language ModelsYuchen Yan, Jin Jiang, Zhenbang Ren, Yijun Li 等ICLR 2026 · 被引用 18 次
- LongCoT: Benchmarking Long-Horizon Chain-of-Thought ReasoningSumeet Motwani, Daniel Nichols, Charles London, Peggy Li 等ICML 2026 · 被引用 2 次
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language ModelsYulei Qin, Gang Li, Zongyi Li, Zihan Xu 等NeurIPS 2025 · 被引用 17 次
- h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement LearningAlesia Ivanova, Sumeet Motwani, Jack Cai, Phil Torr 等ICML 2026 · 被引用 11 次
- Understanding the Role of Training Data in Test-Time ScalingAdel Javanmard, Baharan Mirzasoleiman, Vahab MirrokniICLR 2026 · 被引用 5 次
