Diffuser: Efficient Transformers with Multi-Hop Attention Diffusion for Long Sequences
Aosong Feng, Irene Li, Yuang Jiang, Rex Ying
摘要
Efficient Transformers have been developed for long sequence modeling, due to their subquadratic memory and time complexity. Sparse Transformer is a popular approach to improving the efficiency of Transformers by restricting self-attention to locations specified by the predefined sparse patterns. However, leveraging sparsity may sacrifice expressiveness compared to full-attention, when important token correlations are multiple hops away. To combine advantages of both the efficiency of sparse transformer and the expressiveness of full-attention Transformer, we propose Diffuser, a new state-of-the-art efficient Transformer. Diffuser incorporates all token interactions within one attention layer while maintaining low computation and memory costs. The key idea is to expand the receptive field of sparse attention using Attention Diffusion, which computes multi-hop token correlations based on all paths between corresponding disconnected tokens, besides attention among neighboring tokens. Theoretically, we show the expressiveness of Diffuser as a universal sequence approximator for sequence-to-sequence modeling, and investigate its ability to approximate full-attention by analyzing the graph expander property from the spectral perspective. Experimentally, we investigate the effectiveness of Diffuser with extensive evaluations, including language modeling, image modeling, and Long Range Arena (LRA). Evaluation results show that Diffuser achieves improvements by an average of 0.94% on text classification tasks and 2.30% on LRA, with 1.67x memory savings compared to state-of-the-art benchmarks, which demonstrates superior performance of Diffuser in both expressiveness and efficiency aspects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Rough Transformers: Lightweight and Continuous Time Series Modelling through Signature PatchingFernando Moreno-Pino, Alvaro Arroyo, Harrison Waldon, Xiaowen Dong 等NeurIPS 2024 · 被引用 21 次
- FEASTA: A Flexible and Efficient Accelerator for Sparse Tensor Algebra in Machine LearningKai Zhong, Zhenhua Zhu, Guohao Dai, Hongyi Wang 等ASPLOS 2024 · 被引用 16 次
- AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video GenerationHaoyue Tan, Shengnan Wang, Yulin Qiao, Juncheng Zhang 等CVPR 2026 · 被引用 5 次
- Iterative Sparse Attention for Long-sequence RecommendationGuanyu Lin, Jinwei Luo, Yinfeng Li, Chen Gao 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi 等ICLR 2020 · 被引用 481 次
相关 Paper
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 被引用 11 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- Long-Short Transformer: Efficient Transformers for Language and VisionChen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi 等NeurIPS 2021 · 被引用 180 次
- Combiner: Full Attention Transformer with Sparse Computation CostHongyu Ren, Hanjun Dai, Zihang Dai, Mengjiao Yang 等NeurIPS 2021 · 被引用 99 次
- Trainable Log-linear Sparse Attention for Efficient Diffusion TransformersYifan Zhou, Zeqi Xiao, Tianyi Wei, Shuai Yang 等CVPR 2026 · 被引用 6 次
