Multi Resolution Analysis (MRA) for Approximate Self-Attention
Zhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn Moo Fung, Vikas Singh
Abstract
Transformers have emerged as a preferred model for many tasks in natural langugage processing and vision. Recent efforts on training and deploying Transformers more efficiently have identified many strategies to approximate the self-attention matrix, a key module in a Transformer architecture. Effective ideas include various prespecified sparsity patterns, low-rank basis expansions and combinations thereof. In this paper, we revisit classical Multiresolution Analysis (MRA) concepts such as Wavelets, whose potential value in this setting remains underexplored thus far. We show that simple approximations based on empirical feedback and design choices informed by modern hardware and implementation challenges, eventually yield a MRA-based approach for self-attention with an excellent performance profile across most criteria of interest. We undertake an extensive set of experiments and demonstrate that this multi-resolution scheme outperforms most efficient self-attention proposals and is favorable for both short and long sequences. Code is available at https://github.com/ mlpen/mra-attention .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 878b963e-0117-4ee0-931e-cde080dfb8d2Cited by top-tier papers10
- Radial Attention: 𝒪(n log n) Sparse Attention with Energy Decay for Long Video GenerationXingyang Li, Muyang Li, Tianle Cai, Haocheng Xi et al.NeurIPS 2025 · 66 citations
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing et al.ICLR 2026 · 41 citations
- LookupFFN: Making Transformers Compute-lite for CPU inferenceZhanpeng Zeng, Michael Davies, Pranav Pulijala, Karthikeyan Sankaralingam et al.ICML 2023 · 12 citations
- VCC: Scaling Transformers to 128K Tokens or More by Prioritizing Important TokensZhanpeng Zeng, Cole Hawkins, Mingyi Hong, Aston Zhang et al.NeurIPS 2023 · 11 citations
- Memory Caching: RNNs with Growing MemoryAli Behrouz, Zeman Li, Yuan Deng, Peilin Zhong et al.ICML 2026 · 10 citations
Builds on11
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos et al.ICML 2021 · 1,021 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
Related papers
- Customizing the Inductive Biases of Softmax Attention using Structured MatricesYilun Kuang, Noah Amsel, Sanae Lotfi, Shikai Qiu et al.ICML 2025
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 11 citations
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for SequencesZhenhai Zhu, Radu SoricutACL 2021
- Long-Short Transformer: Efficient Transformers for Language and VisionChen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi et al.NeurIPS 2021 · 180 citations
- Sequence Modeling with Multiresolution Convolutional MemoryJiaxin Shi, Ke Alexander Wang, Emily B. FoxICML 2023 · 24 citations
