BFCL Audio: An Audio Function Calling Evaluation for Large Language Models
Huanzhi Mao, Aditya Ghai, Imra Dawoodani, Tony Ginart, Shishir G. Patil, John Emmons, Joseph E Gonzalez
摘要
Audio agents are increasingly deployed to execute tools from spoken requests, yet audio tool use poses challenges beyond text-only function calling: perception errors (e.g., homophones, noise, disfluencies) can corrupt entities and arguments, and natural interactions often require clarification that changes the tool-calling protocol. We introduce BFCL Audio, a large-scale benchmark for audio function calling with 6.2K expert-verified tasks across two suites that mirror common deployments: BFCL Text Audio (pipelined via transcripts) and BFCL True Audio (end-to-end ). BFCL Audio includes controlled speech and acoustic perturbations (accent and speaking-rate variation, content disfluencies, and background noise) generated through a controllable audio synthesis/augmentation pipeline. We provide automatic grading for both function names and argument values using AST-based matching for single-turn calls and response/state-based metrics for multi-turn interactions, enabling scalable evaluation without LLM judges. Across a broad set of models, we propose a failure-mode taxonomy and analyze which speech and noise factors most strongly impact tool-calling accuracy. We release the benchmark, evaluation harness, and audio pipeline to support research on reliable speech-based agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- -Bench: Evaluating Conversational Agents in a Dual-Control EnvironmentVictor Barres, Honghua Dong, Soham Ray, Xujie Si 等ICML 2026 · 被引用 399 次
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsMinghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song 等EMNLP 2023 · 被引用 72 次
- The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language ModelsShishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji 等ICML 2025
相关 Paper
- HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language ModelsFeiyu Zhao, Yiming Chen, Wenhuan Lu, Daipeng Zhang 等ACL 2026 · 被引用 3 次
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 被引用 5 次
- Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language ModelsZheng Luo, Thirulogasankar Pranav Kutralingam, Ogochukwu N. Okoani, Wanpeng Xu 等ACL 2026 · 被引用 3 次
- GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-CallingHao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang 等ACL 2026
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas 等ICLR 2026 · 被引用 1 次
