RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal Analysis
Enzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong, Qicheng Li, Yong Qin
摘要
Recent advances in speech large language models (e.g., have enabled end-to-end spoken interactions, yet their robustness in realworld applications remains unclear, where systems must assist users in completing specific tasks under complex conditions such as multiturn, ambiguous, and often spontaneous speech, as well as natural alternation between speech and text. Task-oriented dialogue (TOD) offers a realistic scenario to evaluate whether models can effectively help users accomplish such task-oriented goals, but existing benchmarks are mainly text-based, and the few speech datasets are limited to English and often neglect spontaneous disfluencies and speaker diversity. To address this gap, we introduce RealTalk-CN, the first Chinese multi-turn, multi-domain speech-text TOD dataset, containing 5.4k dialogues (60K turns, 150 hours) of real human-to-human recordings with detailed annotations for dialogue states, disfluency types, and speaker characteristics. Based on this dataset, we propose a cross-modal interaction task supporting dynamic speech-text switching and a comprehensive evaluation protocol assessing robustness to disfluencies, sensitivity to speaker variation, and cross-domain generalization. Experiments on state-of-the-art models demonstrate the challenges posed by RealTalk-CN and establish its value as a benchmark for developing reliable and fair Speech LLMs in real-world deployments. The dataset and evaluation framework are available 1 to encourage further research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
- SNIPS: Solving Noisy Inverse Problems StochasticallyBahjat Kawar, Gregory Vaksman, Michael EladNeurIPS 2021 · 被引用 263 次
- GPT-Critic: Offline Reinforcement Learning for End-to-End Task-Oriented Dialogue SystemsYoungsoo Jang, Jongmin Lee, Kee-Eung KimICLR 2022 · 被引用 45 次
- RiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue ModelingJun Quan, Shian Zhang, Qian Cao, Zizhong Li 等EMNLP 2020 · 被引用 41 次
相关 Paper
- TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer CapabilitiesMing Zhang, Caishuang Huang, Yilong Wu, Shichun Liu 等EMNLP 2024
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex ConversationsChengqian Ma, Wei Tao, Steven Y. GuoEMNLP 2025 · 被引用 7 次
- NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven ConversationXiaoyang Wang, Chen Li, Jianqiao Zhao, Dong YuAAAI 2021 · 被引用 54 次
- FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human FeedbackYouquan Li, Miao Zheng, Fan Yang, Guosheng Dong 等EMNLP 2025
- CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog EvaluationYinpei Dai, Wanwei He, Bowen Li, Yuchuan Wu 等EMNLP 2022 · 被引用 6 次
