BFCL Audio: An Audio Function Calling Evaluation for Large Language Models
Huanzhi Mao, Aditya Ghai, Imra Dawoodani, Tony Ginart, Shishir G. Patil, John Emmons, Joseph E Gonzalez
Abstract
Audio agents are increasingly deployed to execute tools from spoken requests, yet audio tool use poses challenges beyond text-only function calling: perception errors (e.g., homophones, noise, disfluencies) can corrupt entities and arguments, and natural interactions often require clarification that changes the tool-calling protocol. We introduce BFCL Audio, a large-scale benchmark for audio function calling with 6.2K expert-verified tasks across two suites that mirror common deployments: BFCL Text Audio (pipelined via transcripts) and BFCL True Audio (end-to-end ). BFCL Audio includes controlled speech and acoustic perturbations (accent and speaking-rate variation, content disfluencies, and background noise) generated through a controllable audio synthesis/augmentation pipeline. We provide automatic grading for both function names and argument values using AST-based matching for single-turn calls and response/state-based metrics for multi-turn interactions, enabling scalable evaluation without LLM judges. Across a broad set of models, we propose a failure-mode taxonomy and analyze which speech and noise factors most strongly impact tool-calling accuracy. We release the benchmark, evaluation harness, and audio pipeline to support research on reliable speech-based agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be836a42-9278-4907-b0f7-89d6a97e3c0dBuilds on4
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- -Bench: Evaluating Conversational Agents in a Dual-Control EnvironmentVictor Barres, Honghua Dong, Soham Ray, Xujie Si et al.ICML 2026 · 399 citations
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsMinghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song et al.EMNLP 2023 · 72 citations
- The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language ModelsShishir G. Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji et al.ICML 2025
Related papers
- HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language ModelsFeiyu Zhao, Yiming Chen, Wenhuan Lu, Daipeng Zhang et al.ACL 2026 · 3 citations
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 5 citations
- Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language ModelsZheng Luo, Thirulogasankar Pranav Kutralingam, Ogochukwu N. Okoani, Wanpeng Xu et al.ACL 2026 · 3 citations
- GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-CallingHao-Xiang Xu, Chong Deng, Jiaqing Liu, Wen Wang et al.ACL 2026
- A Benchmark for Deep Information SynthesisDebjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas et al.ICLR 2026 · 1 citation
