DialogBERT: Discourse-Aware Response Generation via Learning to Recover and Rank Utterances
Xiaodong Gu, Kang Min Yoo, Jung-Woo Ha
Abstract
Recent advances in pre-trained language models have significantly improved neural response generation. However, existing methods usually view the dialogue context as a linear sequence of tokens and learn to generate the next word through token-level self-attention. Such token-level encoding hinders the exploration of discourse-level coherence among utterances. This paper presents DialogBERT, a novel conversational response generation model that enhances previous PLM-based dialogue models. DialogBERT employs a hierarchical Transformer architecture. To efficiently capture the discourse-level coherence among utterances, we propose two training objectives, including masked utterance regression and distributed utterance order ranking in analogy to the original BERT training. Experiments on three multi-turn conversation datasets show that our approach remarkably outperforms three baselines, such as BART and DialoGPT, in terms of quantitative evaluation. The human evaluation suggests that DialogBERT generates more coherent, informative, and human-like responses than the baselines with significant margins.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a91764f-6f88-4493-b9af-bc8ce9aad72cCited by top-tier papers8
- Teacher Forcing Recovers Reward Functions for Text GenerationYongchang Hao, Yuxin Liu, Lili MouNeurIPS 2022 · 24 citations
- Back to the Future: Bidirectional Information Decoupling Network for Multi-turn Dialogue ModelingYiyang Li, Hai Zhao, Zhuosheng ZhangEMNLP 2022 · 9 citations
- KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationQi Zhu, Fei Mi, Zheng Zhang, Yasheng Wang et al.AAAI 2023 · 5 citations
- Smoothing Dialogue States for Open Conversational Machine ReadingZhuosheng Zhang, Siru Ouyang, Hai Zhao, Masao Utiyama et al.EMNLP 2021 · 4 citations
- Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent StructureXueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi et al.EMNLP 2022 · 3 citations
Builds on4
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng et al.AAAI 2020 · 885 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- A Pre-Training Based Personalized Dialogue Generation Model with Persona-Sparse DataYinhe Zheng, Rongsheng Zhang, Minlie Huang, Xiaoxi MaoAAAI 2020 · 173 citations
- Deep Attentive Ranking Networks for Learning to Order SentencesPawan Kumar, Dhanajit Brahma, Harish Karnick, Piyush RaiAAAI 2020 · 52 citations
Related papers
- SLM: Learning a Discourse Language Representation with Sentence UnshufflingHaejun Lee, Drew A. Hudson, Kangwook Lee, Christopher D. ManningEMNLP 2020 · 2 citations
- Long Text Generation by Modeling Sentence-Level and Discourse-Level CoherenceJian Guan, Xiaoxi Mao, Changjie Fan, Zitao Liu et al.ACL 2021
- VD-BERT: A Unified Vision and Dialog Transformer with BERTYue Wang, Shafiq R. Joty, Michael R. Lyu, Irwin King et al.EMNLP 2020 · 68 citations
- Do Response Selection Models Really Know What's Next? Utterance Manipulation Strategies for Multi-turn Response SelectionTaesun Whang, Dongyub Lee, Dongsuk Oh, Chanhee Lee et al.AAAI 2021 · 70 citations
- Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn DialogueLongxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou et al.AAAI 2021 · 57 citations
