Investigating Efficiently Extending Transformers for Long Input Summarization
Jason Phang, Yao Zhao, Peter J. Liu
Abstract
While large pretrained Transformer models have proven highly capable at tackling natural language tasks, handling long sequence inputs still poses a significant challenge. One such task is long input summarization, where inputs are longer than the maximum input context of most models. Through an extensive set of experiments, we investigate what model architectural changes and pretraining paradigms most efficiently adapt a pretrained Transformer for long input summarization. We find that a staggered, block-local Transformer with global encoder tokens strikes a good balance of performance and efficiency, and that an additional pretraining phase on long sequences meaningfully improves downstream summarization performance. Based on our findings, we introduce PEGASUS-X, an extension of the PE-GASUS model with additional long input pretraining to handle inputs of up to 16K tokens, which achieves strong performance on long input summarization tasks comparable with much larger models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a2dafa7-28d2-477d-a221-9c05adcd14fbCited by top-tier papers3
- L2MAC: Large Language Model Automatic Computer for Extensive Code GenerationSamuel Holt, Max Ruiz Luyten, Mihaela van der SchaarICLR 2024 · 29 citations
- DIONYSUS: A Pre-trained Model for Low-Resource Dialogue SummarizationYu Li, Baolin Peng, Pengcheng He, Michel Galley et al.ACL 2023 · 4 citations
- A Sentiment Consolidation Framework for Meta-Review GenerationMiao Li, Jey Han Lau, Eduard H. HovyACL 2024 · 3 citations
Builds on10
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
Related papers
- Long-Span Summarization via Local Attention and Content SelectionPotsawee Manakul, Mark J. F. GalesACL 2021
- Do Long-Range Language Models Actually Use Long-Range Context?Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit IyyerEMNLP 2021 · 35 citations
- SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and DocumentsYusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu et al.ACL 2022
- CoMeT: Collaborative Memory Transformer for Efficient Long Context ModelingRunsong Zhao, Shilei Liu, Jiwei Tang, Langming Liu et al.ACL 2026 · 7 citations
- VCC: Scaling Transformers to 128K Tokens or More by Prioritizing Important TokensZhanpeng Zeng, Cole Hawkins, Mingyi Hong, Aston Zhang et al.NeurIPS 2023 · 11 citations
