R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
Aijia Cheng, Kailong Wang, Ling Shi, Yongxin Zhao
Abstract
Function calling empowers large language models (LLMs) to interface with external tools, yet existing RL-based approaches suffer from misalignment between reasoning processes and tool-call decisions. We propose R2IF, a reasoning-aware RL framework for interpretable function calling, adopting a composite reward integrating format/correctness constraints, Chain-of-Thought Effectiveness Reward (CER), and Specification-Modification-Value (SMV) reward, optimized via GRPO. Experiments on BFCL/ACEBench show R2IF outperforms baselines by up to 34.62% (Llama3.2-3B on BFCL) with positive Average CoT Effectiveness (0.05 for Llama3.2-3B), enhancing both function-calling accuracy and interpretability for reliable tool-augmented LLM deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c76b6ed5-82f8-4a08-8a1f-54ac4436d6f3Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang et al.NeurIPS 2025 · 387 citations
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced ReasoningShaokun Zhang, Yi Dong, Jieyu Zhang, Jan Kautz et al.ICLR 2026 · 61 citations
Related papers
- Improving Large Language Models Function Calling and Interpretability via Guided-Structured TemplatesHy Dang, Tianyi Liu, Zhuofeng Wu, Jingfeng Yang et al.EMNLP 2025
- Empowering LLM Tool Invocation with Tool-call Reward ModelDa Ma, Ziyue Yang, Hongshen Xu, Haotian Fang et al.ICLR 2026
- PROBE: Dense Process Rewards with Observation Evidence for Tool-Augmented Visual ReasoningZongsheng Cao, Anran Liu, Jun Xie, Feng Chen et al.KDD 2026
- CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process SupervisionYifei Lu, Fanghua Ye, Jian Li, Qiang Gao et al.ACL 2025 · 8 citations
- Tool Learning in the Wild: Empowering Language Models as Automatic Tool AgentsZhengliang Shi, Shen Gao, Lingyong Yan, Yue Feng et al.WWW 2025 · 59 citations
