START: Self-taught Reasoner with Tools
Chengpeng Li, Mingfeng Xue, Zhenru Zhang, Jiaxi Yang, Beichen Zhang, Bowen Yu, Binyuan Hui, Junyang Lin, Xiang Wang, Dayiheng Liu
Abstract
Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in complex reasoning through long chain-of-thought, yet they struggle with precise computations and algorithmic operations. Integrating computational tools with LRMs remains challenging, particularly in activating and enhancing models' tooluse capabilities without compromising their reasoning strengths. We address these challenges through START (Self-taught Reasoner with Tools), introducing two key innovations: (1) Hint-infer, a training-free approach that activates LRMs' latent tool-use capabilities through artificial hints, enabling test-time performance scaling; (2) Hint-RFT, a self-training framework that enables models to learn effective tool utilization through diverse hint patterns and rejection-based data synthesis. Experiments show that START significantly improves state-of-the-art LRMs across challenging benchmarks, including competitionlevel mathematics (AMC23: 95.0%, AIME24: 75.6%) and graduate-level science questions (GPQA: 64.6%). Our analysis reveals that START not only enhances accuracy but also improves reasoning efficiency through strategic tool utilization, demonstrating broad applicability in complex reasoning scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72231186-3b5b-48c9-9c36-ce8f9c1cd1b6Cited by top-tier papers5
- Distilling LLM Agent into Small Models with Retrieval and Code ToolsMinki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho et al.NeurIPS 2025 · 51 citations
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference LearningYifei Chen, Guanting Dong, Zhicheng DouICLR 2026 · 18 citations
- Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool CallsZeyu Zhang, Guohao Li, Zhenchang Xing, Alexandros Apostolopoulos et al.ICML 2026 · 1 citation
- Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout TreesKun Li, Zenan Xu, Junan Li, Zengrui Jin et al.ACL 2026 · 1 citation
- Towards Effective Code-Integrated ReasoningFei Bai, Yingqian Min, Beichen Zhang, Zhipeng Chen et al.AAAI 2026
Builds on2
Related papers
- Teaching Language Models to Reason with ToolsChengpeng Li, Zhengyang Tang, Ziniu Li, Mingfeng Xue et al.NeurIPS 2025 · 8 citations
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang et al.ICLR 2026 · 406 citations
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented AgentHaipeng Luo, Huawen Feng, Qingfeng Sun, Can Xu et al.ICLR 2026 · 22 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- AdaReason: Progressive Training of Multi-LoRA Adapters for Budget-Adaptive Language Reasoning ModelsJiacheng Wang, Tianle Chen, Pengyu Cheng, Xiaofeng Hou et al.AAAI 2026
