DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization
Ziming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang, Rui Zhang, Tao Yu, Budhaditya Deb, Chenguang Zhu, Ahmed Hassan Awadallah, Dragomir R. Radev
Abstract
Transformer-based models have achieved state-of-the-art performance on short-input summarization. However, they still struggle with summarizing longer text. In this paper, we present DYLE, a novel dynamic latent extraction approach for abstractive long-input summarization. DYLE jointly trains an extractor and a generator and treats the extracted text snippets as the latent variable, allowing dynamic snippet-level attention weights during decoding. To provide adequate supervision, we propose simple yet effective heuristics for oracle extraction as well as a consistency loss term, which encourages the extractor to approximate the averaged dynamic weights predicted by the generator. We evaluate our method on different long-document and long-dialogue summarization tasks: Gov-Report, QMSum, and arXiv. Experiment results show that DYLE outperforms all existing methods on GovReport and QMSum, with gains up to 6.1 ROUGE, while yielding strong results on arXiv. Further analysis shows that the proposed dynamic weights provide interpretability of our generation process. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a720b9ac-8bf8-4391-a896-0deb93a63cd0Cited by top-tier papers12
- SNaC: Coherence Error Detection for Narrative SummarizationTanya Goyal, Junyi Jessy Li, Greg DurrettEMNLP 2022 · 19 citations
- Leveraging Locality in Abstractive Text SummarizationYixin Liu, Ansong Ni, Linyong Nan, Budhaditya Deb et al.EMNLP 2022 · 19 citations
- How Far are We from Robust Long Abstractive Summarization?Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu et al.EMNLP 2022 · 16 citations
- Factorizing Content and Budget Decisions in Abstractive Summarization of Long DocumentsMarcio Fonseca, Yftah Ziser, Shay B. CohenEMNLP 2022 · 14 citations
- Generating EDU Extracts for Plan-Guided Summary Re-RankingGriffin Adams, Alexander R. Fabbri, Faisal Ladhak, Noémie Elhadad et al.ACL 2023 · 8 citations
Builds on5
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
Related papers
- On Extractive and Abstractive Neural Document Summarization with Transformer Language ModelsJonathan Pilault, Raymond Li, Sandeep Subramanian, Chris PalEMNLP 2020 · 186 citations
- Preserve Context Information for Extract-Generate Long-Input Summarization FrameworkRuifeng Yuan, Zili Wang, Ziqiang Cao, Wenjie LiAAAI 2023 · 3 citations
- Long-Span Summarization via Local Attention and Content SelectionPotsawee Manakul, Mark J. F. GalesACL 2021
- SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and DocumentsYusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu et al.ACL 2022
- Keyword-aware Abstractive Summarization by Extracting Set-level Intermediate SummariesYizhu Liu, Qi Jia, Kenny Q. ZhuWWW 2021 · 14 citations
