Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Shaokun Zhang, Yi Dong, Jieyu Zhang, Jan Kautz, Bryan Catanzaro, Andrew Tao, Qingyun Wu, Zhiding Yu, Guilin Liu
Abstract
Enabling large language models with external tools has become a pivotal strategy for extending their functionality beyond text space. To enhance LLMs' tool-calling abilities, previous approaches primarily rely on supervised fine-tuning (SFT) with trajectories distilled from stronger models, often resulting in imitative reasoning that limits generalization (Chen et al., 2025). In this work, we explore rule-based reinforcement learning (Guo et al., 2025) to enhance tool-calling in LLMs, resulting in Nemotron-Research-Tool-N1, a series of tool-calling reasoning models. Rather than enforcing supervision over intermediate distilled reasoning traces, Tool-N1 1 is trained with a binary RL reward that assesses only the format validity and functional correctness of tool invocations. This lightweight supervision allows the model to develop reasoning strategies independently, without relying on annotated trajectories. Experiments on several major benchmarks show that Tool-N1-7B/14B clearly outperform GPT-4o. We conduct a systematic study on the design of rule-based reinforcement learning strategies for training tool-calling models. Using 5,518 distilled reasoning trajectories, we compare SFT, RL, and the SFT-then-RL pipeline, finding that the widely adopted SFT-then-RL paradigm does not necessarily outperform pure RL. We will release the code in https://github.com/NVlabs/Tool-N1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ee8b6fa-b0a6-4bb1-bb11-01d74e2f64d8Cited by top-tier papers14
- WebDancer: Towards Autonomous Information Seeking AgencyJialong Wu, Baixuan Li, Runnan Fang, Wenbiao Yin et al.NeurIPS 2025 · 194 citations
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic WeightingWenhao Zhang, Yuexiang Xie, Yuchang Sun, Yanxi Chen et al.ICLR 2026 · 100 citations
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool OrchestrationHongjin SU, Shizhe Diao, Ximing Lu, Mingjie Liu et al.ICML 2026 · 34 citations
- Beyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningLiang Chen, Xueting Han, Li Shen, Jing Bai et al.ICML 2026 · 24 citations
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RLSiyi Chen, Mikaela Angelina Uy, Chan Hee Song, Faisal Ladhak et al.CVPR 2026 · 24 citations
Builds on14
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Language Models can Solve Computer TasksGeunwoo Kim, Pierre Baldi, Stephen McAleerNeurIPS 2023 · 539 citations
- Executable Code Actions Elicit Better LLM AgentsXingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang et al.ICML 2024 · 436 citations
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang et al.ICLR 2026 · 406 citations
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
Related papers
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
- Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem SolvingXinji Mai, Haotian Xu, Xing W, Weinong Wang et al.NeurIPS 2025 · 7 citations
- ToolBox-RL: Learning to Generalize Tool Use Across Massive RepositoriesXinyan Shi, Renzhi Wang, Haodong Liu, Piji LiWWW 2026
- AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL SynergyZihan Liu, Zhuolin Yang, Yang Chen, Chankyu Lee et al.ICLR 2026 · 73 citations
- RLP: Reinforcement as a Pretraining ObjectiveAli Hatamizadeh, Syeda Nahida Akter, Shrimai Prabhumoye, Jan Kautz et al.ICLR 2026 · 26 citations
