T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
Zehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang, Dahua Lin, Kai Chen, Feng Zhao
Abstract
Large language models (LLMs) have achieved remarkable performance on various NLP tasks and are augmented by tools for broader applications. Yet, how to evaluate and analyze the tool utilization capability of LLMs is still under-explored. In contrast to previous works that evaluate models holistically, we comprehensively decompose the tool utilization into multiple sub-processes, including instruction following, planning, reasoning, retrieval, understanding, and review. Based on that, we further introduce T-Eval to evaluate the tool-utilization capability step by step. T-Eval disentangles the tool utilization evaluation into several sub-domains along model capabilities, facilitating the inner understanding of both holistic and isolated competency of LLMs. We conduct extensive experiments on T-Eval and in-depth analysis of various LLMs. T-Eval not only exhibits consistency with the outcome-oriented evaluation but also provides a more fine-grained analysis of the capabilities of LLMs, providing a new perspective in LLM evaluation on tool-utilization ability. The benchmark will be available at https://github.com/open-compass/T-Eval .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81c8b363-37bd-4a05-b778-172c244ede08Cited by top-tier papers10
- GUARDIAN: Safeguarding LLM Multi-Agent Collaborations with Temporal Graph ModelingJialong Zhou, Lichao Wang, Xiao YangNeurIPS 2025 · 40 citations
- Benchmarking LLM Tool-Use in the WildPeijie Yu, Wei Liu, Yifan Yang, Jinjian Li et al.ICLR 2026 · 20 citations
- Quality Matters: Evaluating Synthetic Data for Tool-Using LLMsShadi Iskander, Sofia Tolmach, Ori Shapira, Nachshon Cohen et al.EMNLP 2024 · 2 citations
- Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the ModelsJonggeun Lee, Woojung Song, Jongwook Han, Haesung Pyun et al.ACL 2026 · 2 citations
- WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language ModelsShengda Fan, Xin Cong, Yuepeng Fu, Zhong Zhang et al.ICLR 2025 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to UseYue Huang, Jiawen Shi, Yuan Li, Chenrui Fan et al.ICLR 2024 · 188 citations
- FAC²E: Better Understanding Large Language Model Capabilities by Dissociating Language and CognitionXiaoqiang Wang, Lingfei Wu, Tengfei Ma, Bang LiuEMNLP 2024 · 1 citation
- F-Eval: Asssessing Fundamental Abilities with Refined Evaluation MethodsYu Sun, Keyuchen Keyuchen, Shujie Wang, Peiji Li et al.ACL 2024
- TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool UsePengfei He, Zhenwei Dai, Bing He, Hui Liu et al.ICLR 2026 · 46 citations
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language ModelsPei Wang, Yanan Wu, Noah Wang, Jiaheng Liu et al.ICLR 2025
