Lune

ICLR2025

MambaExtend: A Training-Free Approach to Improve Long Context Extension of Mamba

Seyedarmin Azizi, Souvik Kundu, Mohammad Erfan Sadeghi, Massoud Pedram

2025年份

摘要

The inherent quadratic complexity of the attention mechanism in transformer models has driven the research community to explore alternative architectures with sub-quadratic complexity, such as state-space models. Mamba has established itself as a leading model within this emerging paradigm, achieving stateof-the-art results in various language modeling benchmarks. However, despite its impressive performance, Mamba's effectiveness is limited by its pre-training context length, resulting in a pronounced degradation when the model is tasked with handling longer contexts. Our investigation reveals that Mamba's inability to generalize effectively to long contexts is primarily due to the out-of-distribution (OOD) discretization steps. To address this critical limitation, we introduce Mam-baExtend, a novel framework designed to significantly enhance the context extension capabilities of Mamba. Specifically, MambaExtend leverages a training-free approach to calibrate only the scaling factors of discretization modules for different layers. We demonstrate both gradient-based and gradient-free zeroth-order optimization to learn the optimal scaling factors for each Mamba layer, requiring orders of magnitude fewer updates as opposed to the parameter fine-tuning-based alternatives. Using this approach, we achieve a training-free context extension of up to 32×, expanding the context from 2k to 64k tokens with minimal increases in perplexity. In contrast to existing fine-tuning methods, MambaExtend selectively calibrates the scaling factors, requiring up to ∼5.42 * 10 6 × fewer parameter updates and incurring up to 3.87× lower peak memory usage, while delivering comparable or superior long-context performance across multiple tasks. Codes and checkpoints are available here 1 .

Published as a conference paper at ICLR 2025 different approach to handling long sequences at sub-quadratic complexity. Unlike transformers, SSMs are grounded in continuous-time dynamics and offer the potential to handle much longer sequences without blowing out the memory and compute demand. Mamba (Gu & Dao, 2023; Dao & Gu, 2024), a popular SSM variant built leveraging the selective state-space layers (S6), has shown impressive performance on various NLP, image, and medical genomics benchmarks (Schiff et al., 2024). The key advantage of Mamba stems from the sub-quadratic compute complexity of theoretically grounded linear RNN layers.

LLMs for long-context understanding have recently found many useful applications, including summarizing long documents and answering long questions (Chen et al., 2023c). However, transformerbased LLMs that are pre-trained on fixed-length contexts yield lower generative performance when used on longer sequences during inference time (Chen et al., 2024; 2023b). This shortcoming of the transformers is tied to the inability of the positional embedding to generalize well on longer sequences, causing such sequences to appear as out-of-distribution (OOD) sequences (Chen et al., 2023c;Jin et al., 2024). Interestingly, Mamba models, despite their theoretical ability to capture global interactions, also fail to generalize to long sequence or context lengths (Ben-Kish et al., 2024). This phenomenon has been tied to the Mamba model's implicit bias to a limited effective receptive field (ERF) governed by the training data sequence length (Ben-Kish et al., 2024).

For transformer-based LLMs, the OOD sequence length generalization has been explored extensively, including fine-tuning to longer sequences (Chen et al., 2023c) and allowing sophisticated modification to the transformer's positional embedding (Jin et al., 2024;Ding et al., 2024;Golovneva et al., 2024). Unfortunately, such solutions are not directly applicable to Mamba models. This is primarily due to the absence of an explicit positional embedding for Mamba models to generalize. Moreover, unlike transformers, the potential root cause of Mamba's performance deterioration for long sequence processing is yet to be discovered.