Jamba: Hybrid Transformer-Mamba Language Models
Barak Lenz, Opher Lieber, Alan Arazi, Amir Bergman, Avshalom Manevich, Barak Peleg, Ben Aviram, Chen Almagor, Clara Fridman, Dan Padnos, Daniel Gissin, Daniel Jannai
Abstract
We present Jamba, a new base large language model based on a novel hybrid Transformer-Mamba mixture-of-experts (MoE) architecture. Specifically, Jamba interleaves blocks of Transformer and Mamba layers, enjoying the benefits of both model families. MoE is added in some of these layers to increase model capacity while keeping active parameter usage manageable. This flexible architecture allows resource-and objective-specific configurations. In the particular configuration we have implemented, we end up with a powerful model that fits in a single 80GB GPU. Built at large scale, Jamba provides high throughput and small memory footprint compared to vanilla Transformers, and at the same time state-of-the-art performance on standard language model benchmarks and long-context evaluations. Remarkably, the model presents strong results for up to 256K tokens context length. We study various architectural decisions, such as how to combine Transformer and Mamba layers, and how to mix experts, and show that some of them are crucial in large scale modeling. We also describe several interesting properties of these architectures which the training and evaluation of Jamba have revealed, and plan to release checkpoints from various ablation runs, to encourage further exploration of this novel architecture. We make the weights of our implementation of Jamba publicly available under a permissive license. Model: https://huggingface.co/ai21labs/Jamba-v0.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 288174b5-d2f9-40ce-a7e4-0cc79c3f2e69Cited by top-tier papers12
- HAMLET: Switch Your Vision-Language-Action Model into a History-Aware PolicyMyungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee et al.ICLR 2026 · 52 citations
- TabSTAR: A Tabular Foundation Model for Tabular Data with Text FieldsAlan Arazi, Eilam Shapira, Roi ReichartNeurIPS 2025 · 20 citations
- Generalizable, real-time neural decoding with hybrid state-space modelsAvery Hee-Woon Ryoo, Nanda H. Krishna, Ximeng Mao, Mehdi Azabou et al.NeurIPS 2025 · 16 citations
- tttLRM: Test-Time Training for Long Context and Autoregressive 3D ReconstructionChen Wang, Hao Tan, Wang Yifan, Zhiqin Chen et al.CVPR 2026 · 11 citations
- Overcoming Long Context Limitations of State Space Models via Context Dependent Sparse AttentionZhihao Zhan, Jianan Zhao, Zhaocheng Zhu, Jian TangNeurIPS 2025 · 7 citations
Builds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language ModelYixing Li, Ruobing Xie, Zhen Yang, Xingwu Sun et al.AAAI 2026 · 3 citations
- B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading MemoryLuca Zancato, Arjun Seshadri, Yonatan Dukler, Aditya Golatkar et al.NeurIPS 2024 · 34 citations
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen et al.ICLR 2025
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky et al.ICML 2026
- Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence ModelingSili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang et al.NeurIPS 2024 · 65 citations
