Think Less, Act Warranted: Efficient Tool-Integrated Reasoning via Dual-Efficiency Regularization
Yichen Xiao, Siyu Gong, Linan Yue
Abstract
Recent methods using Reinforcement Learning (RL) have improved Tool-Integrated Reasoning (TIR) by training large language models to learn end-to-end policies for multi-step tool usage, enabling them to solve complex tasks more effectively. Despite these advances, existing methods often suffer from overthinking at both the action and reasoning levels: models tend to invoke tools redundantly and generate excessively long reasoning trajectories, resulting in high computational cost. To address this, in this paper, we propose LightTIR, a dual-penalty reward framework, to achieve efficient TIR. For action efficiency, LightTIR estimates the marginal utility of each tool call through prefix-aligned counterfactual trajectories, encouraging calls that contribute meaningful information while penalizing low-utility or redundant invocations. For reasoning efficiency, LightTIR introduces a length-aware regularization term, adaptively penalizing intermediate reasoning steps that exceed the minimal effective trajectory required for correct prediction. Extensive experiments demonstrate that LightTIR can reduce redundancy and trajectory expansion while maintaining answer correctness, achieving more efficient RL-based TIR. Code is available at https://github.com/ekventitas/LightTIR.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 23ad5dbb-0bd1-463d-8c4c-06fc25a0c9acRelated papers
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference LearningYifei Chen, Guanting Dong, Zhicheng DouICLR 2026 · 18 citations
- Rethinking LLM Reasoning: From Explicit Trajectories to Latent RepresentationsCong Jiang, Xiaofeng Zhang, Fangzhi Zhu, XiaoWei Chen et al.ICLR 2026
- DEPO: Dual-Efficiency Preference Optimization for LLM AgentsSirui Chen, Mengshi Zhao, Lei Xu, Yuying Zhao et al.AAAI 2026 · 2 citations
- MatchTIR: Fine-Grained Supervision for Tool-Integrated Reasoning via Bipartite MatchingChangle Qu, Sunhao Dai, Hengyi Cai, Jun Xu et al.ACL 2026 · 3 citations
- TInR: Exploring Tool-Internalized Reasoning in Large Language ModelsQiancheng Xu, Yongqi Li, Fan Liu, Hongru Wang et al.ACL 2026
