SuperDialseg: A Large-scale Dataset for Supervised Dialogue Segmentation
Junfeng Jiang, Chengzhang Dong, Sadao Kurohashi, Akiko Aizawa
Abstract
Dialogue segmentation is a crucial task for dialogue systems allowing a better understanding of conversational texts. Despite recent progress in unsupervised dialogue segmentation methods, their performances are limited by the lack of explicit supervised signals for training. Furthermore, the precise definition of segmentation points in conversations still remains as a challenging problem, increasing the difficulty of collecting manual annotations. In this paper, we provide a feasible definition of dialogue segmentation points with the help of document-grounded dialogues and release a large-scale supervised dataset called SuperDialseg, containing 9,478 dialogues based on two prevalent document-grounded dialogue corpora, and also inherit their useful dialogue-related annotations. Moreover, we provide a benchmark including 18 models across five categories for the dialogue segmentation task with several proper evaluation metrics. Empirical studies show that supervised learning is extremely effective in in-domain datasets and models trained on SuperDialseg can achieve good generalization ability on out-of-domain data. Additionally, we also conducted human verification on the test set and the Kappa score confirmed the quality of our automatically constructed dataset. We believe our work is an important step forward in the field of dialogue segmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented DialogueChien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming XiongEMNLP 2020 · 210 citations
- DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationMing Zhong, Yang Liu, Yichong Xu, Chenguang Zhu et al.AAAI 2022 · 150 citations
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 121 citations
- Topic-Aware Multi-turn Dialogue ModelingYi Xu, Hai Zhao, Zhuosheng ZhangAAAI 2021 · 93 citations
Related papers
- VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic TransitionsYuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li et al.ACL 2023 · 4 citations
- doc2dial: A Goal-Oriented Document-Grounded Dialogue DatasetSong Feng, Hui Wan, R. Chulaka Gunasekara, Siva Sankalp Patel et al.EMNLP 2020 · 87 citations
- SalesBot: Transitioning from Chit-Chat to Task-Oriented DialoguesSsu Chiu, Maolin Li, Yen-Ting Lin, Yun-Nung ChenACL 2022
- ClidSum: A Benchmark Dataset for Cross-Lingual Dialogue SummarizationJiaan Wang, Fandong Meng, Ziyao Lu, Duo Zheng et al.EMNLP 2022 · 27 citations
- MedDialog: Large-scale Medical Dialogue DatasetsGuangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang et al.EMNLP 2020 · 163 citations
