CodeSSM: Towards State Space Models for Code Understanding
Shweta Verma, Abhinav Anand, Mira Mezini
Abstract
Although transformers dominate many codespecific tasks, they have significant limitations. This paper explores State Space Models (SSMs) as a promising alternative for code understanding tasks such as retrieval, classification, and clone detection. We introduce CodeSSM, the first SSM-based model trained on code corpora to assess its effectiveness. Our results demonstrate that SSMs are more sampleefficient and can extrapolate to longer contexts beyond the pretraining length. Extensive experiments show that SSMs offer a viable alternative to transformers, addressing several their limitations. Additionally, CodeSSM reduces memory usage by up to 64% compared to transformers at a context length of 2048, with greater savings as context length grows. The code is available here. CodeSSM In this section, we introduce the basic architecture of CodeSSM and its variations that we investigate. We also explain the pretraining setup. Architecture CodeSSM is an encoder-only model consisting of 12 layers 1 . This model is built upon the Bidirec-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dee44505-7388-4ff8-beb9-d7f4d607092dBuilds on23
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Combining Recurrent, Convolutional, and Continuous-time Models with Linear State Space LayersAlbert Gu, Isys Johnson, Karan Goel, Khaled Saab et al.NeurIPS 2021 · 1,280 citations
Related papers
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space ModelsEran Malach, Omid Saremi, Sinead Williamson, Arwen Bradley et al.ICLR 2026 · 3 citations
- Repeat After Me: Transformers are Better than State Space Models at CopyingSamy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran MalachICML 2024 · 176 citations
- On the Expressiveness and Length Generalization of Selective State Space Models on Regular LanguagesAleksandar Terzic, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann et al.AAAI 2025 · 8 citations
- Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate MechanismAviv Bick, Eric P. Xing, Albert GuICML 2025
- Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothingPeihao Wang, Ruisi Cai, Yuehao Wang, Jiajun Zhu et al.ICLR 2025
