Structural Pre-training for Dialogue Comprehension
Zhuosheng Zhang, Hai Zhao
Abstract
Pre-trained language models (PrLMs) have demonstrated superior performance due to their strong ability to learn universal language representations from self-supervised pre-training. However, even with the help of the powerful PrLMs, it is still challenging to effectively capture task-related knowledge from dialogue texts which are enriched by correlations among speaker-aware utterances. In this work, we present SPIDER, Structural Pre-traIned DialoguE Reader, to capture dialogue exclusive features. To simulate the dialoguelike features, we propose two training objectives in addition to the original LM objectives: 1) utterance order restoration, which predicts the order of the permuted utterances in dialogue context; 2) sentence backbone regularization, which regularizes the model to improve the factual correctness of summarized subject-verb-object triplets. Experimental results on widely used dialogue benchmarks verify the effectiveness of the newly introduced self-supervised tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Structural Characterization for Dialogue DisentanglementXinbei Ma, Zhuosheng Zhang, Hai ZhaoACL 2022 · 20 citations
- Dialog-Post: Multi-Level Self-Supervised Objectives and Hierarchical Model for Dialogue Post-TrainingZhenyu Zhang, Lei Shen, Yuming Zhao, Meng Chen et al.ACL 2023 · 3 citations
- Language Model Pre-training on True NegativesZhuosheng Zhang, Hai Zhao, Masao Utiyama, Eiichiro SumitaAAAI 2023 · 3 citations
- Pre-training Multi-party Dialogue Models with Latent Discourse InferenceYiyang Li, Xinting Huang, Wei Bi, Hai ZhaoACL 2023 · 3 citations
Builds on9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu et al.ACL 2020 · 229 citations
- MuTual: A Dataset for Multi-Turn Dialogue ReasoningLeyang Cui, Yu Wu, Shujie Liu, Yue Zhang et al.ACL 2020 · 115 citations
- Topic-Aware Multi-turn Dialogue ModelingYi Xu, Hai Zhao, Zhuosheng ZhangAAAI 2021 · 93 citations
Related papers
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue SystemYixuan Su, Lei Shu, Elman Mansimov, Arshit Gupta et al.ACL 2022 · 218 citations
- Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn DialogueLongxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou et al.AAAI 2021 · 57 citations
- Delving into Global Dialogue Structures: Structure Planning Augmented Response Selection for Multi-turn ConversationsTingchen Fu, Xueliang Zhao, Rui YanKDD 2023 · 8 citations
- FutureTOD: Teaching Future Knowledge to Pre-trained Language Model for Task-Oriented DialogueWeihao Zeng, Keqing He, Yejie Wang, Chen Zeng et al.ACL 2023 · 3 citations
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long et al.EMNLP 2020 · 60 citations
