On Extractive and Abstractive Neural Document Summarization with Transformer Language Models
Jonathan Pilault, Raymond Li, Sandeep Subramanian, Chris Pal
Abstract
We present a method to produce abstractive summaries of long documents that exceed several thousand words via neural abstractive summarization. We perform a simple extractive step before generating a summary, which is then used to condition the transformer language model on relevant information before being tasked with generating a summary. We also show that this approach produces more abstractive summaries compared to prior work that employs a copy mechanism while still achieving higher ROUGE scores. We provide extensive comparisons with strong baseline methods, prior state of the art work as well as multiple variants of our approach including those using only transformers, only extractive techniques and combinations of the two. We examine these models using four different summarization tasks and datasets: arXiv papers, PubMed papers, the Newsroom and BigPatent datasets. We find that transformer based methods produce summaries with fewer n-gram copies, leading to n-gram copying statistics that are more similar to human generated abstracts. We include a human evaluation, finding that transformers are ranked highly for coherence and fluency, but purely extractive methods score higher for informativeness and relevance. We hope that these architectures and experiments may serve as strong points of comparison for future work. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72a3ba78-331d-43fa-ae0c-23d20533e813Cited by top-tier papers38
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- In-Context Impersonation Reveals Large Language Models' Strengths and BiasesLeonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz et al.NeurIPS 2023 · 259 citations
- Lift Yourself Up: Retrieval-augmented Text Generation with Self-MemoryXin Cheng, Di Luo, Xiuying Chen, Lemao Liu et al.NeurIPS 2023 · 177 citations
- Poolingformer: Long Document Modeling with Pooling AttentionHang Zhang, Yeyun Gong, Yelong Shen, Weisheng Li et al.ICML 2021 · 118 citations
Builds on1
Related papers
- Long-Span Summarization via Local Attention and Content SelectionPotsawee Manakul, Mark J. F. GalesACL 2021
- DYLE: Dynamic Latent Extraction for Abstractive Long-Input SummarizationZiming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang et al.ACL 2022 · 62 citations
- SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and DocumentsYusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu et al.ACL 2022
- Enriching and Controlling Global Semantics for Text SummarizationThong Nguyen, Anh Tuan Luu, Truc Lu, Tho QuanEMNLP 2021 · 26 citations
- Semantic Self-Segmentation for Abstractive Summarization of Long Documents in Low-Resource RegimesGianluca Moro, Luca RagazziAAAI 2022 · 67 citations
