-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
Soham Ray, Keshav Dhandhania, Victor Barres, Karthik Narasimhan
Abstract
Full-duplex voice agents-systems that listen and speak simultaneously-are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We introduce τ -Voice, a benchmark for evaluating voice agents on grounded tasks with real-world complexity: agents must navigate complex multi-turn conversations, adhere to domain policies, and interact with the environment. The framework extends τ 2bench into a novel voice agent benchmark combining verifiable completion of complex grounded tasks, full-duplex interaction, and realistic audio-enabling direct comparison between voice and text performance. A controllable and realistic voice user simulator provides diverse accents, realistic audio environments, and rich turn-taking dynamics; by decoupling simulation from wallclock time, the user simulator can use the most capable LLM without real-time constraints. We evaluate task completion (pass@1) and voice interaction quality across 278 tasks: while GPT-5 (reasoning) achieves 85%, voice agents reach only 31-51% under clean conditions and 26-38% under realistic conditions with noise and diverse accents-retaining only 30-45% of text capability; qualitative analysis confirms 79-90% of failures stem from agent behavior, suggesting that observed failures primarily reflect agent behavior under our evaluation setup. τ -Voice provides a reproducible testbed for measuring progress toward voice agents that are natural, conversational, and reliable.
- Equal contribution , listed in reverse alphabetical order.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7585db78-2bb2-4983-94c4-928969314a1bBuilds on8
- -Bench: Evaluating Conversational Agents in a Dual-Control EnvironmentVictor Barres, Honghua Dong, Soham Ray, Xujie Si et al.ICML 2026 · 399 citations
- SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex ConversationWenyi Yu, Siyin Wang, Xiaoyu Yang, Xianzhao Chen et al.NeurIPS 2025 · 43 citations
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech InteractionShu-Wen Yang, Ming Tu, Ting-Wei Liu, Xinghua Qu et al.ICLR 2026 · 29 citations
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech ModelsFeng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue et al.ACL 2026 · 17 citations
- Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human InteractionAdvait Gosai, Tyler Vuong, Utkarsh Tyagi, Steven Li et al.ACL 2026 · 10 citations
Related papers
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsShunyu Yao, Noah Shinn, Pedram Razavi, Karthik R. NarasimhanICLR 2025
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous EnvironmentsRomain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja et al.ICLR 2026 · 29 citations
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?Liqiang Jing, Zhehui Huang, Xiaoyang Wang, Wenlin Yao et al.ICLR 2025
- CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World UncertaintyJohannes Kirmayr, Lukas Stappen, Elisabeth AndréACL 2026 · 5 citations
- GitTaskBench: A Benchmark for Code Agents Solving Real-World Tasks Through Code Repository LeveragingZiyi Ni, Huacan Wang, Shuo Zhang, Shuo Lu et al.AAAI 2026 · 13 citations
