Overcoming Long Context Limitations of State Space Models via Context Dependent Sparse Attention
Zhihao Zhan, Jianan Zhao, Zhaocheng Zhu, Jian Tang
摘要
Efficient long-context modeling remains a critical challenge for natural language processing (NLP), as the time complexity of the predominant Transformer architecture scales quadratically with the sequence length. While state-space models (SSMs) offer alternative sub-quadratic solutions, they struggle to capture longrange dependencies effectively. In this work, we focus on analyzing and improving the long-context modeling capabilities of SSMs. We show that the widely used synthetic task, associative recall, which requires a model to recall a value associated with a single key without context, insufficiently represents the complexities of real-world long-context modeling. To address this limitation, we extend the associative recall to a novel synthetic task, joint recall, which requires a model to recall the value associated with a key given in a specified context. Theoretically, we prove that SSMs do not have the expressiveness to solve multi-query joint recall in sub-quadratic time complexity. To resolve this issue, we propose a solution based on integrating SSMs with Context-Dependent Sparse Attention (CDSA), which has the expressiveness to solve multi-query joint recall with sub-quadratic computation. To bridge the gap between theoretical analysis and real-world applications, we propose locality-sensitive Hashing Attention with sparse Key Selection (HAX), which instantiates the theoretical solution and is further tailored to natural language domains. Extensive experiments on both synthetic and real-world long-context benchmarks show that HAX consistently outperforms SSM baselines and SSMs integrated with context-independent sparse attention (CISA). Our code is available at: https://github.com/DeepGraphLearning/HAX.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand 等ICLR 2026 · 被引用 11 次
- Expressivity-Efficiency Tradeoffs for Hybrid Sequence ModelsJohn Cooper, Mingchen Ma, Ilias Diakonikolas, Frederic SalaICML 2026
它引用的顶会 Paper21
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra 等NeurIPS 2020 · 被引用 1,100 次
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 等ICML 2023 · 被引用 481 次
相关 Paper
- Sparse Attention with Learning to HashZhiqing Sun, Yiming Yang, Shinjae YooICLR 2022 · 被引用 21 次
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language ModelsXiang Hu, Zhanchao Zhou, Ruiqi Liang, Zehuan Li 等ACL 2026 · 被引用 2 次
- Scaling Linear Attention Capacity with Sparse State ExpansionYuqi Pan, Yongqi An, Zheng Li, Yuhong Chou 等ICLR 2026 · 被引用 3 次
- Linear-Time Self Attention with Codeword Histogram for Efficient RecommendationYongji Wu, Defu Lian, Neil Zhenqiang Gong, Lu Yin 等WWW 2021 · 被引用 18 次
- MATCH: Modulating Attention via In-Context Retrieval for Long-Context TransformersLinrui Ma, Chun Hei Lo, Xinyu Wang, Peng Lu 等ACL 2026
