Towards Effective Code-Integrated Reasoning
Fei Bai, Yingqian Min, Beichen Zhang, Zhipeng Chen, Xin Zhao, Lei Fang, Zheng Liu, Zhongyuan Wang, Hongteng Xu
摘要
In this paper, we investigate code-integrated reasoning (CIR), where models generate code when necessary and integrate feedback by executing it through a code interpreter. To acquire this capability, models must learn when and how to use external code tools effectively, which is supported by tool-augmented reinforcement learning (RL). Despite its benefits, tool-augmented RL can still suffer from potential instability in the learning dynamics. In light of this challenge, we present a systematic approach ETIR (Effective TIR) to improving the training effectiveness and stability of tool-augmented RL for code-integrated reasoning. Specifically, we develop enhanced training strategies that balance exploration and stability, progressively building tool-use capabilities while improving reasoning performance. Through extensive experiments on five mainstream mathematical reasoning benchmarks, our model demonstrates significant performance improvements over multiple competitive baselines. Furthermore, we conduct an in-depth analysis of the mechanism of code-integrated reasoning, revealing several key insights, such as the extension of model’s capability boundaries and the simultaneous improvement of reasoning efficiency through code integration. These findings underscore the potential of code-integrated reasoning as a scalable paradigm for advancing robust and efficient language model reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated ReasoningZhenghai Xue, Longtao Zheng, Qian Liu, Yingru Li 等ICLR 2026 · 被引用 152 次
- Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement LearningRan Xu, Jingjing Chen, Jiayu Ye, Yu Wu 等ICLR 2026 · 被引用 17 次
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement LearningYulei Qin, Xiaoyu Tan, Zhengbao He, Gang Li 等ICLR 2026 · 被引用 9 次
- How Many Code and Test Cases Are Enough? Evaluating Test Cases Generation from a Binary-Matrix PerspectiveXianzhen Luo, Jinyang Huang, Wenzhen Zheng, Qingfu Zhu 等ICLR 2026 · 被引用 5 次
- Learning from the Irrecoverable: Error-Localized Policy Optimization for Tool-Integrated LLM ReasoningQiao Liang, Yuke Zhu, Chao Ge, Lei Yang 等ACL 2026 · 被引用 4 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang 等NeurIPS 2025 · 被引用 1,109 次
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang 等ICLR 2026 · 被引用 406 次
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem SolvingZhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen 等ICLR 2024 · 被引用 289 次
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai 等AAAI 2026 · 被引用 216 次
相关 Paper
- Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem SolvingXinji Mai, Haotian Xu, Xing W, Weinong Wang 等NeurIPS 2025 · 被引用 7 次
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical ReasoningQikai Chang, Zhenrong Zhang, Pengfei Hu, Jun Du 等ICLR 2026 · 被引用 8 次
- TInR: Exploring Tool-Internalized Reasoning in Large Language ModelsQiancheng Xu, Yongqi Li, Fan Liu, Hongru Wang 等ACL 2026
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference LearningYifei Chen, Guanting Dong, Zhicheng DouICLR 2026 · 被引用 18 次
- Think Less, Act Warranted: Efficient Tool-Integrated Reasoning via Dual-Efficiency RegularizationYichen Xiao, Siyu Gong, Linan YueKDD 2026
