RiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue Modeling
Jun Quan, Shian Zhang, Qian Cao, Zizhong Li, Deyi Xiong
摘要
In order to alleviate the shortage of multidomain data and to capture discourse phenomena for task-oriented dialogue modeling, we propose RiSAWOZ, a large-scale multidomain Chinese Wizard-of-Oz dataset with Rich Semantic Annotations. RiSAWOZ contains 11.2K human-to-human (H2H) multiturn semantically annotated dialogues, with more than 150K utterances spanning over 12 domains, which is larger than all previous annotated H2H conversational datasets. Both single-and multi-domain dialogues are constructed, accounting for 65% and 35%, respectively. Each dialogue is labeled with comprehensive dialogue annotations, including dialogue goal in the form of natural language description, domain, dialogue states and acts at both the user and system side. In addition to traditional dialogue annotations, we especially provide linguistic annotations on discourse phenomena, e.g., ellipsis and coreference, in dialogues, which are useful for dialogue coreference and ellipsis resolution tasks. Apart from the fully annotated dataset, we also present a detailed description of the data collection procedure, statistics and analysis of the dataset. A series of benchmark models and results are reported, including natural language understanding (intent detection & slot filling), dialogue state tracking and dialogue contextto-text generation, as well as coreference and ellipsis resolution, which facilitate the baseline comparison for future research on this corpus. 1 * Equal Contributions. 1 The corpus is publicly available at https://github. com/terryqj0107/RiSAWOZ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Fusing Task-Oriented and Open-Domain Dialogues in Conversational AgentsTom Young, Frank Xing, Vlad Pandelea, Jinjie Ni 等AAAI 2022 · 被引用 64 次
- SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence UnderstandingTianyu Yu, Chengyue Jiang, Chao Lou, Shen Huang 等AAAI 2024 · 被引用 30 次
- Towards Measuring and Modeling "Culture" in LLMs: A SurveyMuhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh 等EMNLP 2024 · 被引用 21 次
- Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User GoalsZeming Liu, Jun Xu, Zeyang Lei, Haifeng Wang 等ACL 2022 · 被引用 18 次
- Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text ClassificationJiachen Tian, Shizhan Chen, Xiaowang Zhang, Zhiyong Feng 等EMNLP 2021 · 被引用 13 次
它引用的顶会 Paper2
- Towards Scalable Multi-Domain Conversational Agents: The Schema-Guided Dialogue DatasetAbhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta 等AAAI 2020 · 被引用 707 次
- Filling Conversation Ellipsis for Better Social Dialog UnderstandingXiyuan Zhang, Chengxi Li, Dian Yu, Samuel Davidson 等AAAI 2020 · 被引用 17 次
相关 Paper
- MA-DST: Multi-Attention-Based Scalable Dialog State TrackingAdarsh Kumar, Peter Ku, Anuj Kumar Goyal, Angeliki Metallinou 等AAAI 2020 · 被引用 61 次
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong 等ACL 2026
- Multi-domain Dialogue State Tracking with Recursive InferenceLizi Liao, Tongyao Zhu, Le Hong Long, Tat-Seng ChuaWWW 2021 · 被引用 11 次
- CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog EvaluationYinpei Dai, Wanwei He, Bowen Li, Yuchuan Wu 等EMNLP 2022 · 被引用 6 次
- Multi-Domain Dialogue Acts and Response Co-GenerationKai Wang, Junfeng Tian, Rui Wang, Xiaojun Quan 等ACL 2020 · 被引用 46 次
