NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
Kinjal Basu, Ibrahim Abdelaziz, Kiran Kate, Mayank Agarwal, Maxwell Crouse, Yara Rizk, Kelsey Bradford, Asim Munawar, Sadhana Kumaravel, Saurabh Goyal, Xin Wang, Luis A. Lastras, Pavan Kapanipathi
摘要
The resurgence of autonomous agents built using large language models (LLMs) to solve complex real-world tasks has brought increased focus on LLMs' fundamental ability of tool or function calling. At the core of these agents, an LLM must plan, execute, and respond using external tools, APIs, and custom functions. Research on tool calling has gathered momentum, but evaluation benchmarks and datasets representing the complexity of the tasks have lagged behind. In this work, we focus on one such complexity, nested sequencing, with the goal of extending existing benchmarks and evaluation. Specifically, we present NESTFUL , a benchmark to evaluate LLMs on nested sequences of API calls, i.e., sequences where the output of one API call is passed as input to a subsequent call. NESTFUL contains 1800+ nested sequences where all the function calls are executable. Experimental results on a variety of models show that the bestperforming model (GPT-4o) achieves a full sequence match accuracy of 28% and a winrate of 60%, necessitating a large scope for improvement in the nested sequencing aspect of function calling. Our analysis of these results provides possible future research directions for the community, in addition to a benchmark to track progress. We have released the NEST-FUL dataset under the Apache 2.0 license at https://github.com/IBM/NESTFUL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Magnet: Multi-turn Tool-use Data Synthesis and Distillation via Graph TranslationFan Yin, Zifeng Wang, I-Hung Hsu, Jun Yan 等ACL 2025 · 被引用 23 次
- Procedural Environment Generation for Tool-Use AgentsMichael Sullivan, Mareike Hartmann, Alexander KollerEMNLP 2025 · 被引用 12 次
- LlamaRestTest: Effective REST API Testing with Small Language ModelsMyeongsoo Kim, Saurabh Sinha, Alessandro OrsoFSE 2025 · 被引用 9 次
- A Multi-Agent Approach for REST API Testing with Semantic Graphs and LLM-Driven InputsMyeongsoo Kim, Tyler Stennett, Saurabh Sinha, Alessandro OrsoICSE 2025 · 被引用 4 次
- CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error ScenariosShiting Huang, Zhen Fang, Zehui Chen, Siyu Yuan 等EMNLP 2025
它引用的顶会 Paper7
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMsKinjal Basu, Ibrahim Abdelaziz, Subhajit Chaudhury, Soham Dan 等ACL 2024 · 被引用 6 次
- ShortcutsBench: A Large-Scale Real-world Benchmark for API-based AgentsHaiyang Shen, Yue Li, Desong Meng, Dongqi Cai 等ICLR 2025
相关 Paper
- AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API CallsYu Du, Fangyun Wei, Hongyang ZhangICML 2024 · 被引用 104 次
- The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language ModelsShishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji 等ICML 2025
- AppBench: Planning of Multiple APIs from Various APPs for Complex User InstructionHongru Wang, Rui Wang, Boyang Xue, Heming Xia 等EMNLP 2024 · 被引用 2 次
- CONFETTI: Conversational Function-Calling Evaluation Through Turn-Level InteractionsTamer Alkhouli, Katerina Margatina, James Gung, Raphael Shu 等ACL 2025
- Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long ContextsYifei Yu, Qian-Wen Zhang, Lingfeng Qiao, Di Yin 等EMNLP 2025 · 被引用 2 次
