VoxMind: An End-to-End Agentic Spoken Dialogue System
Tianle Liang, Yifu Chen, Shengpeng Ji, Yijun Chen, Zhiyang Jia, Jingyu Lu, Fan Zhuo, Xueyi Pu, Yangzhuo Li, Zhou Zhao
摘要
Recent end-to-end spoken dialogue models enable natural interaction. However, as user demands become increasingly complex, models that rely solely on conversational abilities often struggle to cope. Incorporating agentic capabilities is therefore essential: by enabling tool use, these models can extend their knowledge boundaries and better solve real-world tasks. Yet, existing research has largely concentrated on core perception and generation, with comparatively limited exploration of such tool-augmented extensions. To bridge this gap, we present VoxMind, an integrated framework designed to equip end-to-end spoken dialogue models with comprehensive agentic abilities. Leveraging our curated 470-hour AgentChat dataset, we incorporate a "Thinkbefore-Speak" mechanism, enabling the model to internalize structured reasoning as a critical prerequisite for planning and response generation. Furthermore, to mitigate latency bottlenecks caused by large-scale tool integration, we propose a Multi-Agent Dynamic Tool Management architecture. By asynchronously delegating retrieval tasks to an auxiliary agent aligned with the main model's reasoning trajectory, this system effectively decouples inference latency from toolset size. Experimental results confirm that VoxMind achieves significant improvements in agent performance: compared with strong baselines, the task completion rate increases from 34.88% to 74.57%, outperforming Gemini-2.5-Pro on spoken agent tasks while preserving general conversational quality. The source code and associated data are publicly available at https: //github.com/MM-Speech/VoxMind .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool UsageSiddhant Arora, Haidar Khan, Kai Sun, Xin Dong 等ICML 2026 · 被引用 22 次
- ReAct: Synergizing Reasoning and Acting in Language ModelsShunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du 等ICLR 2023
相关 Paper
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic ToolsJunde Wu, Jiayuan Zhu, Yuyuan Liu, Min Xu 等ACL 2025 · 被引用 88 次
- ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded ExecutionShouzheng Huang, Meishan Zhang, Baotian Hu, Min ZhangACL 2026
- AutoTool: Dynamic Tool Selection and Integration for Agentic ReasoningJiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen 等ICML 2026 · 被引用 4 次
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use MixtureYongchao Chen, Jiefeng Chen, Rui Meng, Ji Yin 等ICLR 2026 · 被引用 13 次
- ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue AgentsZhigen Li, Jianxiang Peng, Yanmeng Wang, Yong Cao 等ACL 2025 · 被引用 9 次
