Poolingformer: Long Document Modeling with Pooling Attention
Hang Zhang, Yeyun Gong, Yelong Shen, Weisheng Li, Jiancheng Lv, Nan Duan, Weizhu Chen
Abstract
In this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling. Its first level uses a smaller sliding window pattern to aggregate information from neighbors. Its second level employs a larger window to increase receptive fields with pooling attention to reduce both computational cost and memory consumption. We first evaluate Poolingformer on two long sequence QA tasks: the monolingual NQ and the multilingual TyDi QA. Experimental results show that Poolingformer sits atop three official leaderboards measured by F1, outperforming previous state-of-the-art models by 1.9 points (79.8 vs. 77.9) on NQ long answer, 1.9 points (79.5 vs. 77.6) on TyDi QA passage answer, and 1.6 points (67.6 vs. 66.0) on TyDi QA minimal answer. We further evaluate Poolingformer on a long sequence summarization task. Experimental results on the arXiv benchmark continue to demonstrate its superior performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1efb3a41-41fd-4fb8-83f9-9ed752b7431dCited by top-tier papers19
- Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge GraphJiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang et al.ICLR 2024 · 247 citations
- DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationMing Zhong, Yang Liu, Yichong Xu, Chenguang Zhu et al.AAAI 2022 · 150 citations
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu et al.NeurIPS 2022 · 104 citations
- Paths-over-Graph: Knowledge Graph Empowered Large Language Model ReasoningXingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu et al.WWW 2025 · 86 citations
- Sparse Binary Transformers for Multivariate Time Series ModelingMatt Gorbett, Hossein Shirazi, Indrakshi RayKDD 2023 · 18 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
Related papers
- Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM EmbeddingsXueying Ding, Xingyue Huang, Mingxuan Ju, Liam Collins et al.ACL 2026 · 3 citations
- ResFormer: All-Time Reservoir Memory for Long Sequence ClassificationHongbo Liu, Jia XuEMNLP 2025
- RikiNet: Reading Wikipedia Pages for Natural Question AnsweringDayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan et al.ACL 2020 · 55 citations
- Document Modeling with Graph Attention Networks for Multi-grained Machine Reading ComprehensionBo Zheng, Haoyang Wen, Yaobo Liang, Nan Duan et al.ACL 2020 · 57 citations
- SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and DocumentsYusen Zhang, Ansong Ni, Ziming Mao, Chen Henry Wu et al.ACL 2022
