DialogLM: Pre-trained Model for Long Dialogue Understanding and Summarization
Ming Zhong, Yang Liu, Yichong Xu, Chenguang Zhu, Michael Zeng
摘要
Dialogue is an essential part of human communication and cooperation. Existing research mainly focuses on short dialogue scenarios in a one-on-one fashion. However, multi-person interactions in the real world, such as meetings or interviews, are frequently over a few thousand words. There is still a lack of corresponding research and powerful tools to understand and process such long dialogues. Therefore, in this work, we present a pre-training framework for long dialogue understanding and summarization. Considering the nature of long conversations, we propose a window-based denoising approach for generative pre-training. For a dialogue, it corrupts a window of text with dialogue-inspired noise, and guides the model to reconstruct this window based on the content of the remaining conversation. Furthermore, to process longer input, we augment the model with sparse attention which is combined with conventional attention in a hybrid manner. We conduct extensive experiments on five datasets of long dialogues, covering tasks of dialogue summarization, abstractive question answering and topic segmentation. Experimentally, we show that our pre-trained model DialogLM significantly surpasses the state-of-the-art models across datasets and tasks. Source code and all the pre-trained models are available on our GitHub repository (https://github.com/microsoft/DialogLM).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- ClidSum: A Benchmark Dataset for Cross-Lingual Dialogue SummarizationJiaan Wang, Fandong Meng, Ziyao Lu, Duo Zheng 等EMNLP 2022 · 被引用 27 次
- MeetingBank: A Benchmark Dataset for Meeting SummarizationYebowen Hu, Timothy Ganter, Hanieh Deilamsalehy, Franck Dernoncourt 等ACL 2023 · 被引用 19 次
- MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic DialoguesKuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu 等AAAI 2025 · 被引用 10 次
- ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue SummarizationXiutian Zhao, Ke Wang, Wei PengEMNLP 2023 · 被引用 5 次
- DIONYSUS: A Pre-trained Model for Low-Resource Dialogue SummarizationYu Li, Baolin Peng, Pengcheng He, Michel Galley 等ACL 2023 · 被引用 4 次
它引用的顶会 Paper15
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
相关 Paper
- SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and DocumentsYusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu 等ACL 2022
- Preserve Context Information for Extract-Generate Long-Input Summarization FrameworkRuifeng Yuan, Zili Wang, Ziqiang Cao, Wenjie LiAAAI 2023 · 被引用 3 次
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu 等ACL 2020 · 被引用 229 次
- Language Model as an Annotator: Exploring DialoGPT for Dialogue SummarizationXiachong Feng, Xiaocheng Feng, Libo Qin, Bing Qin 等ACL 2021
- Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph DenoiseZhenghao Lin, Yeyun Gong, Yelong Shen, Tong Wu 等ICML 2023 · 被引用 107 次
