Sequence Modeling with Multiresolution Convolutional Memory
Jiaxin Shi, Ke Alexander Wang, Emily B. Fox
Abstract
Efficiently capturing the long-range patterns in sequential data sources salient to a given task -- such as classification and generative modeling -- poses a fundamental challenge. Popular approaches in the space tradeoff between the memory burden of brute-force enumeration and comparison, as in transformers, the computational burden of complicated sequential dependencies, as in recurrent neural networks, or the parameter burden of convolutional networks with many or large filters. We instead take inspiration from wavelet-based multiresolution analysis to define a new building block for sequence modeling, which we call a MultiresLayer. The key component of our model is the multiresolution convolution, capturing multiscale trends in the input sequence. Our MultiresConv can be implemented with shared filters across a dilated causal convolution tree. Thus it garners the computational advantages of convolutional networks and the principled theoretical motivation of wavelet decompositions. Our MultiresLayer is straightforward to implement, requires significantly fewer parameters, and maintains at most a memory footprint for a length sequence. Yet, by stacking such layers, our model yields state-of-the-art performance on a number of sequence classification and autoregressive density estimation tasks using CIFAR-10, ListOps, and PTB-XL datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60859523-c8d4-4589-98ca-181cbd3d0185Cited by top-tier papers11
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing et al.ICLR 2026 · 41 citations
- Parallelizing non-linear sequential models over the sequence lengthYi Heng Lim, Qi Zhu, Joshua Selfridge, Muhammad Firmansyah KasimICLR 2024 · 33 citations
- Hybrid2 Neural ODE Causal Modeling and an Application to Glycemic ResponseBob Junyi Zou, Matthew E. Levine, Dessi P. Zaharieva, Ramesh Johari et al.ICML 2024 · 13 citations
- Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long SequencesZicheng Liu, Siyuan Li, Li Wang, Zedong Wang et al.ICML 2024 · 11 citations
- Recursion in Recursion: Two-Level Nested Recursion for Length Generalization with ScalabilityJishnu Ray Chowdhury, Cornelia CarageaNeurIPS 2023 · 7 citations
Builds on17
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
Related papers
- Reparameterized Multi-Resolution Convolutions for Long Sequence ModellingJake Cunningham, Giorgio Giannone, Mingtian Zhang, Marc Peter DeisenrothNeurIPS 2024 · 4 citations
- Approximation Theory of Convolutional Architectures for Time Series ModellingHaotian Jiang, Zhong Li, Qianxiao LiICML 2021 · 14 citations
- Multi Resolution Analysis (MRA) for Approximate Self-AttentionZhanpeng Zeng, Sourav Pal, Jeffery Kline, Glenn Moo Fung et al.ICML 2022 · 14 citations
- MEGABYTE: Predicting Million-byte Sequences with Multiscale TransformersLili Yu, Daniel Simig, Colin Flaherty, Armen Aghajanyan et al.NeurIPS 2023 · 197 citations
- MELODI: Exploring Memory Compression for Long ContextsYinpeng Chen, DeLesley Hutchins, Aren Jansen, Andrey Zhmoginov et al.ICLR 2025 · 1 citation
