Long-Span Summarization via Local Attention and Content Selection
Potsawee Manakul, Mark J. F. Gales
摘要
Transformer-based models have achieved state-of-the-art results in a wide range of natural language processing (NLP) tasks including document summarization. Typically these systems are trained by fine-tuning a large pretrained model to the target task. One issue with these transformer-based models is that they do not scale well in terms of memory and compute requirements as the input length grows. Thus, for long document summarization, it can be challenging to train or fine-tune these models. In this work, we exploit large pre-trained transformer-based models and address long-span dependencies in abstractive summarization using two methods: local self-attention; and explicit content selection. These approaches are compared on a range of network configurations. Experiments are carried out on standard long-span summarization tasks, including Spotify Podcast, arXiv, and PubMed datasets. We demonstrate that by combining these methods, we can achieve state-of-the-art results on all three tasks in the ROUGE scores. Moreover, without a large-scale GPU card, our approach can achieve comparable or better results than existing approaches. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Salience Allocation as Guidance for Abstractive SummarizationFei Wang, Kaiqiang Song, Hongming Zhang, Lifeng Jin 等EMNLP 2022 · 被引用 26 次
- Leveraging Locality in Abstractive Text SummarizationYixin Liu, Ansong Ni, Linyong Nan, Budhaditya Deb 等EMNLP 2022 · 被引用 19 次
- How Far are We from Robust Long Abstractive Summarization?Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu 等EMNLP 2022 · 被引用 16 次
- Discourse-Aware Soft Prompting for Text GenerationMarjan Ghazvininejad, Vladimir Karpukhin, Vera Gor, Asli CelikyilmazEMNLP 2022 · 被引用 6 次
- ResearchPulse: Building Method-Experiment Chains through Multi-Document Scientific InferenceQi Chen, Jingxuan Wei, Zhuoya Yao, Haiguang Wang 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
相关 Paper
- On Extractive and Abstractive Neural Document Summarization with Transformer Language ModelsJonathan Pilault, Raymond Li, Sandeep Subramanian, Chris PalEMNLP 2020 · 被引用 186 次
- Semantic Self-Segmentation for Abstractive Summarization of Long Documents in Low-Resource RegimesGianluca Moro, Luca RagazziAAAI 2022 · 被引用 67 次
- DYLE: Dynamic Latent Extraction for Abstractive Long-Input SummarizationZiming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang 等ACL 2022 · 被引用 62 次
- Investigating Efficiently Extending Transformers for Long Input SummarizationJason Phang, Yao Zhao, Peter J. LiuEMNLP 2023 · 被引用 30 次
- SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and DocumentsYusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu 等ACL 2022
