Pre-training Multi-party Dialogue Models with Latent Discourse Inference
Yiyang Li, Xinting Huang, Wei Bi, Hai Zhao
Abstract
Multi-party dialogues are more difficult for models to understand than one-to-one two-party dialogues, since they involve multiple interlocutors, resulting in interweaving reply-to relations and information flows. To step over these obstacles, an effective way is to pre-train a model that understands the discourse structure of multi-party dialogues, namely, to whom each utterance is replying. However, due to the lack of explicitly annotated discourse labels in multi-party dialogue corpora, previous works fail to scale up the pre-training process by putting aside the unlabeled multi-party conversational data for nothing. To fully utilize the unlabeled data, we propose to treat the discourse structures as latent variables, then jointly infer them and pre-train the discourse-aware model by unsupervised latent variable inference methods. Experiments on multiple downstream tasks show that our pre-trained model outperforms strong baselines by large margins and achieves state-of-the-art (SOTA) results, justifying the effectiveness of our method. The official implementation of this paper is available at https://github.com/EricLee8/MPD_EMVI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbe35d9f-185e-423b-a3ad-20370f4634c4Builds on12
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu et al.ACL 2020 · 229 citations
- Multi-turn Response Selection using Dialogue Dependency RelationsQi Jia, Yizhu Liu, Siyu Ren, Kenny Q. Zhu et al.EMNLP 2020 · 31 citations
- Structural Characterization for Dialogue DisentanglementXinbei Ma, Zhuosheng Zhang, Hai ZhaoACL 2022 · 20 citations
- Back to the Future: Bidirectional Information Decoupling Network for Multi-turn Dialogue ModelingYiyang Li, Hai Zhao, Zhuosheng ZhangEMNLP 2022 · 9 citations
Related papers
- EM Pre-training for Multi-party Dialogue Response GenerationYiyang Li, Hai ZhaoACL 2023 · 9 citations
- MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation UnderstandingJia-Chen Gu, Chongyang Tao, Zhen-Hua Ling, Can Xu et al.ACL 2021
- Is Discourse Role Important for Emotion Recognition in Conversation?Donovan Ong, Jian Su, Bin Chen, Anh Tuan Luu et al.AAAI 2022 · 29 citations
- Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent StructureXueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi et al.EMNLP 2022 · 3 citations
- Masking Orchestration: Multi-Task Pretraining for Multi-Role Dialogue Representation LearningTianyi Wang, Yating Zhang, Xiaozhong Liu, Changlong Sun et al.AAAI 2020 · 8 citations
