MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
Yingyue Li, Bencheng Liao, Wenyu Liu, Xinggang Wang
摘要
With the advancement of RNN models with linear complexity, the quadratic complexity challenge of transformers has the potential to be overcome. Notably, the emerging Mamba-2 has demonstrated competitive performance, bridging the gap between RNN models and transformers. However, due to sequential processing and vanishing gradients, RNN models struggle to capture long-range dependencies, limiting contextual understanding. This results in slow convergence, high resource demands, and poor performance on downstream understanding and complex reasoning tasks. In this work, we present MaTVLM, a method for distilling pre-trained vision-language models (VLMs) into an efficient Mamba-Transformer hybrid architecture. Specifically, we construct this hybrid architecture by replacing a portion of the transformer decoder layers in the pre-trained VLM with Mamba-2 layers. Building on this design, we employ a single-stage distillation process, incorporating a clever initialization strategy, leveraging the inherent relationship between attention mechanisms and Mamba-2, and initialize Mamba-2 with corresponding attention weights, which notably accelerates convergence. With the pre-trained VLM serving as the teacher model, this distillation process further boosts both convergence speed and model performance. Furthermore, we investigate the impact of differential distillation loss within our training framework. We evaluate MaTVLM on multiple benchmarks, demonstrating competitive performance against the teacher model and existing VLMs while surpassing both Mamba-based VLMs and models of comparable parameter scales. Remarkably, MaTVLM achieves up to faster inference than the teacher model while reducing GPU memory consumption by 27.5%, all without compromising performance. Code and models are released at https://github.com/hustvl/MaTVLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video UnderstandingBoshen Xu, Zihan Xiao, Jiaze Li, Jianzhong Ju 等CVPR 2026 · 被引用 5 次
- Scaling Parallel Sequence Models to Vision Foundation ModelsYitong Jiang, Collin McCarthy, Hongjun Wang, Hanrong Ye 等CVPR 2026
它引用的顶会 Paper31
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
相关 Paper
- ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision ModelsGuoyizhe Wei, Rama ChellappaICCV 2025
- The Mamba in the Llama: Distilling and Accelerating Hybrid ModelsJunxiong Wang, Daniele Paliotta, Avner May, Alexander M. Rush 等NeurIPS 2024 · 被引用 146 次
- VoCo-LLaMA: Towards Vision Compression with Large Language ModelsXubing Ye, Yukang Gan, Xiaoke Huang, Yixiao Ge 等CVPR 2025
- VAMBA: Understanding Hour-Long Videos with Hybrid Mamba-TransformersWeiming Ren, Wentao Ma, Huan Yang, Cong Wei 等ICCV 2025 · 被引用 2 次
- MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive PretrainingYunze Liu, Li YiCVPR 2025
