Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
Jiarui Jiang, Wei Huang, Miao Zhang, Taiji Suzuki, Liqiang Nie
Abstract
State-space models (SSMs), particularly Mamba, emerge as an efficient Transformer alternative with linear complexity for long-sequence modeling. Recent empirical works demonstrate Mamba's in-context learning (ICL) capabilities competitive with Transformers, a critical capacity for large foundation models. However, theoretical understanding of Mamba's ICL remains limited, restricting deeper insights into its underlying mechanisms. Even fundamental tasks such as linear regression ICL, widely studied as a standard theoretical benchmark for Transformers, have not been thoroughly analyzed in the context of Mamba. To address this gap, we study the training dynamics of Mamba on the linear regression ICL task. By developing novel techniques tackling non-convex optimization with gradient descent related to Mamba's structure, we establish an exponential convergence rate to ICL solution, and derive a loss bound that is comparable to Transformer's. Importantly, our results reveal that Mamba can perform a variant of online gradient descent to learn the latent function in context. This mechanism is different from that of Transformer, which is typically understood to achieve ICL through gradient descent emulation. The theoretical results are verified by experimental simulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda et al.ICML 2024 · 390 citations
Related papers
- Can Mamba Learn How To Learn? A Comparative Study on In-Context Learning TasksJongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee et al.ICML 2024 · 124 citations
- From Markov to Laplace: How Mamba In-Context Learns Markov ChainsMarco Bondaschi, Nived Rajaraman, Xiuying Wei, Razvan Pascanu et al.ICLR 2026 · 10 citations
- Longhorn: State Space Models are Amortized Online LearnersBo Liu, Rui Wang, Lemeng Wu, Yihao Feng et al.ICLR 2025
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic dataTianyi Chen, Pengxiao Lin, Zhiwei Wang, Zhi-Qin John XuNeurIPS 2025 · 4 citations
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language ModelYixing Li, Ruobing Xie, Zhen Yang, Xingwu Sun et al.AAAI 2026 · 3 citations
