Jamba: Hybrid Transformer-Mamba Language Models
Barak Lenz, Opher Lieber, Alan Arazi, Amir Bergman, Avshalom Manevich, Barak Peleg, Ben Aviram, Chen Almagor, Clara Fridman, Dan Padnos, Daniel Gissin, Daniel Jannai
摘要
We present Jamba, a new base large language model based on a novel hybrid Transformer-Mamba mixture-of-experts (MoE) architecture. Specifically, Jamba interleaves blocks of Transformer and Mamba layers, enjoying the benefits of both model families. MoE is added in some of these layers to increase model capacity while keeping active parameter usage manageable. This flexible architecture allows resource-and objective-specific configurations. In the particular configuration we have implemented, we end up with a powerful model that fits in a single 80GB GPU. Built at large scale, Jamba provides high throughput and small memory footprint compared to vanilla Transformers, and at the same time state-of-the-art performance on standard language model benchmarks and long-context evaluations. Remarkably, the model presents strong results for up to 256K tokens context length. We study various architectural decisions, such as how to combine Transformer and Mamba layers, and how to mix experts, and show that some of them are crucial in large scale modeling. We also describe several interesting properties of these architectures which the training and evaluation of Jamba have revealed, and plan to release checkpoints from various ablation runs, to encourage further exploration of this novel architecture. We make the weights of our implementation of Jamba publicly available under a permissive license. Model: https://huggingface.co/ai21labs/Jamba-v0.1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- HAMLET: Switch Your Vision-Language-Action Model into a History-Aware PolicyMyungkyu Koo, Daewon Choi, Taeyoung Kim, Kyungmin Lee 等ICLR 2026 · 被引用 52 次
- TabSTAR: A Tabular Foundation Model for Tabular Data with Text FieldsAlan Arazi, Eilam Shapira, Roi ReichartNeurIPS 2025 · 被引用 20 次
- Generalizable, real-time neural decoding with hybrid state-space modelsAvery Hee-Woon Ryoo, Nanda H. Krishna, Ximeng Mao, Mehdi Azabou 等NeurIPS 2025 · 被引用 16 次
- tttLRM: Test-Time Training for Long Context and Autoregressive 3D ReconstructionChen Wang, Hao Tan, Wang Yifan, Zhiqin Chen 等CVPR 2026 · 被引用 11 次
- Overcoming Long Context Limitations of State Space Models via Context Dependent Sparse AttentionZhihao Zhan, Jianan Zhao, Zhaocheng Zhu, Jian TangNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- TransMamba: A Sequence-Level Hybrid Transformer-Mamba Language ModelYixing Li, Ruobing Xie, Zhen Yang, Xingwu Sun 等AAAI 2026 · 被引用 3 次
- B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading MemoryLuca Zancato, Arjun Seshadri, Yonatan Dukler, Aditya Golatkar 等NeurIPS 2024 · 被引用 34 次
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen 等ICLR 2025
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky 等ICML 2026
- Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence ModelingSili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang 等NeurIPS 2024 · 被引用 65 次
