Sparsifying Transformer Models with Trainable Representation Pooling
Michal Pietruszka, Lukasz Borchmann, Lukasz Garncarek
摘要
We propose a novel method to sparsify attention in the Transformer model by learning to select the most-informative token representations during the training process, thus focusing on the task-specific parts of an input. A reduction of quadratic time and memory complexity to sublinear was achieved due to a robust trainable top-k operator.Our experiments on a challenging long document summarization task show that even our simple baseline performs comparably to the current SOTA, and with trainable pooling we can retain its top quality, while being 1.8faster during training, 4.5faster during inference, and up to 13more computationally efficient in the decoder.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Linear Complexity Randomized Self-attention MechanismLin Zheng, Chong Wang, Lingpeng KongICML 2022 · 被引用 39 次
- Leveraging Locality in Abstractive Text SummarizationYixin Liu, Ansong Ni, Linyong Nan, Budhaditya Deb 等EMNLP 2022 · 被引用 19 次
- LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token MergingZhijian Shu, Cheng Lin, Tao Xie, Wei Yin 等CVPR 2026 · 被引用 17 次
- How Far are We from Robust Long Abstractive Summarization?Huan Yee Koh, Jiaxin Ju, He Zhang, Ming Liu 等EMNLP 2022 · 被引用 16 次
它引用的顶会 Paper8
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li 等ICCV 2021 · 被引用 1,611 次
- Sparse Sinkhorn AttentionYi Tay, Dara Bahri, Liu Yang, Donald Metzler 等ICML 2020 · 被引用 391 次
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingZihang Dai, Guokun Lai, Yiming Yang, Quoc LeNeurIPS 2020 · 被引用 273 次
相关 Paper
- Sparse is Enough in Scaling TransformersSebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser 等NeurIPS 2021 · 被引用 127 次
- Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token SelectionDongwon Jo, Beomseok Kang, Jiwon Song, jae-joon kimICML 2026 · 被引用 1 次
- SEA: Sparse Linear Attention with Estimated Attention MaskHeejun Lee, Jina Kim, Jeffrey Willette, Sung Ju HwangICLR 2024 · 被引用 12 次
- QuoKA: Query-Oriented KV Selection for Efficient LLM PrefillDalton Jones, Junyoung Park, Matthew J. Morse, Mingu Lee 等ICLR 2026 · 被引用 4 次
- Delta Attention: Fast and Accurate Sparse Attention Inference by Delta CorrectionJeffrey Willette, Heejun Lee, Sung Ju HwangNeurIPS 2025 · 被引用 9 次
