InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning
Yuchen Yan, Liang Jiang, Jin Jiang, Shuaicheng Li, zujie wen, Zhiqiang Zhang, JUN ZHOU, Jian Shao, Yueting Zhuang, Yongliang Shen
摘要
Large reasoning models achieve strong performance by scaling inference-time chain-ofthought, but this paradigm suffers from quadratic cost, context length limits, and degraded reasoning due to lost-in-the-middle effects. Iterative reasoning mitigates these issues by periodically summarizing intermediate thoughts, yet existing methods rely on supervised learning or fixed heuristics and fail to optimize when to summarize, what to preserve, and how to resume reasoning. We propose InftyThink + , an end-toend reinforcement learning framework that optimizes the entire iterative reasoning trajectory, building on model-controlled iteration boundaries and explicit summarization. InftyThink + adopts a two-stage training scheme with supervised coldstart followed by trajectory-level reinforcement learning, enabling the model to learn strategic summarization and continuation decisions. Experiments on DeepSeek-R1-Distill-Qwen-1.5B show that InftyThink + improves accuracy by 21% on AIME24 and outperforms conventional long chain-of-thought reinforcement learning by a clear margin, while also generalizing better to out-of-distribution benchmarks. Moreover, InftyThink + significantly reduces inference latency and accelerates reinforcement learning training, demonstrating improved reasoning efficiency alongside stronger performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language ModelsYuchen Yan, Yongliang Shen, Yang Liu, Jin Jiang 等ICLR 2026 · 被引用 48 次
- Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal StructureZirui Li, Xuefeng Bai, Kehai Chen, Yizhi Li 等ICML 2026
- User-Aware Active Knowledge Acquisition for Emotional Support DialogueMufan Xu, Kehai Chen, Jiahao Hu, Xinchao Xu 等ICML 2026
它引用的顶会 Paper5
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- OpenThoughts: Data Recipes for Reasoning ModelsEtash Kumar Guha, Ryan Marten, Sedrick Keh, Negin Raoof 等ICLR 2026 · 被引用 235 次
- Agentic Reinforced Policy OptimizationGuanting Dong, Hangyu Mao, Kai Ma, Licheng Bao 等ICLR 2026 · 被引用 146 次
- Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning ModelsXingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He 等ICML 2025
- The Markovian Thinker: Architecture-Agnostic Linear Scaling of ReasoningMilad Aghajohari, Kamran Chitsaz, Amirhossein Kazemnejad, Sarath Chandar 等ICLR 2026
相关 Paper
- Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RLSongjun Tu, Jiahao Lin, Qichao Zhang, Xiangyu Tian 等NeurIPS 2025 · 被引用 69 次
- Your Models Have Thought Enough: Training Large Reasoning Models to Stop OverthinkingJinyi Han, Ying Huang, Ying Liao, Haiquan Zhao 等ICLR 2026 · 被引用 11 次
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang 等ACL 2026 · 被引用 8 次
- How Far Are We from Optimal Reasoning Efficiency?Jiaxuan Gao, Shu Yan, Qixin Tan, Lu Yang 等NeurIPS 2025 · 被引用 12 次
- Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated PenaltyZewei Yu, Lirong Gao, Yuke Zhu, Bo Zheng 等ICLR 2026 · 被引用 2 次
