RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal Analysis
Enzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong, Qicheng Li, Yong Qin
Abstract
Recent advances in speech large language models (e.g., have enabled end-to-end spoken interactions, yet their robustness in realworld applications remains unclear, where systems must assist users in completing specific tasks under complex conditions such as multiturn, ambiguous, and often spontaneous speech, as well as natural alternation between speech and text. Task-oriented dialogue (TOD) offers a realistic scenario to evaluate whether models can effectively help users accomplish such task-oriented goals, but existing benchmarks are mainly text-based, and the few speech datasets are limited to English and often neglect spontaneous disfluencies and speaker diversity. To address this gap, we introduce RealTalk-CN, the first Chinese multi-turn, multi-domain speech-text TOD dataset, containing 5.4k dialogues (60K turns, 150 hours) of real human-to-human recordings with detailed annotations for dialogue states, disfluency types, and speaker characteristics. Based on this dataset, we propose a cross-modal interaction task supporting dynamic speech-text switching and a comprehensive evaluation protocol assessing robustness to disfluencies, sensitivity to speaker variation, and cross-domain generalization. Experiments on state-of-the-art models demonstrate the challenges posed by RealTalk-CN and establish its value as a benchmark for developing reliable and fair Speech LLMs in real-world deployments. The dataset and evaluation framework are available 1 to encourage further research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 24d7fc30-6512-40e1-842e-0601500d0f54Builds on8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- SNIPS: Solving Noisy Inverse Problems StochasticallyBahjat Kawar, Gregory Vaksman, Michael EladNeurIPS 2021 · 263 citations
- GPT-Critic: Offline Reinforcement Learning for End-to-End Task-Oriented Dialogue SystemsYoungsoo Jang, Jongmin Lee, Kee-Eung KimICLR 2022 · 45 citations
- RiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue ModelingJun Quan, Shian Zhang, Qian Cao, Zizhong Li et al.EMNLP 2020 · 41 citations
Related papers
- TransferTOD: A Generalizable Chinese Multi-Domain Task-Oriented Dialogue System with Transfer CapabilitiesMing Zhang, Caishuang Huang, Yilong Wu, Shichun Liu et al.EMNLP 2024
- C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex ConversationsChengqian Ma, Wei Tao, Steven Y. GuoEMNLP 2025 · 7 citations
- NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven ConversationXiaoyang Wang, Chen Li, Jianqiao Zhao, Dong YuAAAI 2021 · 54 citations
- FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human FeedbackYouquan Li, Miao Zheng, Fan Yang, Guosheng Dong et al.EMNLP 2025
- CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog EvaluationYinpei Dai, Wanwei He, Bowen Li, Yuchuan Wu et al.EMNLP 2022 · 6 citations
