Residual Shuffle-Exchange Networks for Fast Processing of Long Sequences
Andis Draguns, Emils Ozolins, Agris Sostaks, Matiss Apinis, Karlis Freivalds
Abstract
Attention is a commonly used mechanism in sequence processing, but it is of O(n^2) complexity which prevents its application to long sequences. The recently introduced neural Shuffle-Exchange network offers a computation-efficient alternative, enabling the modelling of long-range dependencies in O(n log n) time. The model, however, is quite complex, involving a sophisticated gating mechanism derived from the Gated Recurrent Unit. In this paper, we present a simple and lightweight variant of the Shuffle-Exchange network, which is based on a residual network employing GELU and Layer Normalization. The proposed architecture not only scales to longer sequences but also converges faster and provides better accuracy. It surpasses the Shuffle-Exchange network on the LAMBADA language modelling task and achieves state-of-the-art performance on the MusicNet dataset for music transcription while being efficient in the number of parameters. We show how to combine the improved Shuffle-Exchange network with convolutional layers, establishing it as a useful building block in long sequence processing applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 380d5a90-81d1-495b-b22c-c2faa935ca57Cited by top-tier papers1
- Bayesian Sparsification of Deep C-valued NetworksIvan Nazarov, Evgeny BurnaevICML 2020 · 4 citations
Builds on2
Related papers
- ChordMixer: A Scalable Neural Attention Model for Sequences with Different LengthRuslan Khalitov, Tong Yu, Lei Cheng, Zhirong YangICLR 2023 · 4 citations
- Toeplitz Neural Network for Sequence ModelingZhen Qin, Xiaodong Han, Weixuan Sun, Bowen He et al.ICLR 2023 · 5 citations
- Not All Attention Is Needed: Gated Attention Network for Sequence DataLanqing Xue, Xiaopeng Li, Nevin L. ZhangAAAI 2020 · 47 citations
- Forgetting Transformer: Softmax Attention with a Forget GateZhixuan Lin, Evgenii Nikishin, Xu Owen He, Aaron C. CourvilleICLR 2025
- Luna: Linear Unified Nested AttentionXuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou et al.NeurIPS 2021 · 145 citations
