Sparse Modular Activation for Efficient Sequence Modeling
Liliang Ren, Yang Liu, Shuohang Wang, Yichong Xu, Chenguang Zhu, ChengXiang Zhai
Abstract
Linear State Space Models (SSMs) have demonstrated strong performance in a variety of sequence modeling tasks due to their efficient encoding of the recurrent structure. However, in more comprehensive tasks like language modeling and machine translation, self-attention-based models still outperform SSMs. Hybrid models employing both SSM and self-attention generally show promising performance, but current approaches apply attention modules statically and uniformly to all elements in the input sequences, leading to sub-optimal quality-efficiency trade-offs. In this work, we introduce Sparse Modular Activation (SMA), a general mechanism enabling neural networks to sparsely and dynamically activate sub-modules for sequence elements in a differentiable manner. Through allowing each element to skip non-activated sub-modules, SMA reduces computation and memory consumption at both training and inference stages of sequence modeling. As a specific instantiation of SMA, we design a novel neural architecture, SeqBoat, which employs SMA to sparsely activate a Gated Attention Unit (GAU) based on the state representations learned from an SSM. By constraining the GAU to only conduct local attention on the activated inputs, SeqBoat can achieve linear inference complexity with theoretically infinite attention span, and provide substantially better quality-efficiency trade-off than the chunking-based models. With experiments on a wide range of tasks, including language modeling, speech classification and long-range arena, SeqBoat brings new state-of-the-art results among hybrid models with linear complexity and reveals the amount of attention needed for each task through the learned sparse activation patterns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2da0884d-6d42-4ffd-ba9f-63a82d4a60b6Cited by top-tier papers14
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
- Mechanistic Design and Scaling of Hybrid ArchitecturesMichael Poli, Armin W. Thomas, Eric Nguyen, Pragaash Ponnusamy et al.ICML 2024 · 57 citations
- Selective Attention: Enhancing Transformer through Principled Context ControlXuechen Zhang, Xiangyu Chang, Mingchen Li, Amit K. Roy-Chowdhury et al.NeurIPS 2024 · 32 citations
- State-Free Inference of State-Space Models: The Transfer Function ApproachRom N. Parnichkun, Stefano Massaroli, Alessandro Moro, Jimmy T. H. Smith et al.ICML 2024 · 18 citations
- Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long GenerationLiliang Ren, Congcong Chen, Haoran Xu, Young Jin Kim et al.NeurIPS 2025 · 12 citations
Builds on22
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long SequencesZicheng Liu, Siyuan Li, Li Wang, Zedong Wang et al.ICML 2024 · 11 citations
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen et al.ICLR 2025
- SEA: Sparse Linear Attention with Estimated Attention MaskHeejun Lee, Jina Kim, Jeffrey Willette, Sung Ju HwangICLR 2024 · 12 citations
- Block-State TransformersJonathan Pilault, Mahan Fathi, Orhan Firat, Chris Pal et al.NeurIPS 2023 · 33 citations
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 11 citations
