Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
Yuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen, Shang Yang, Song Han, Han Cai
Abstract
We present Jet-Nemotron, a new family of hybrid-architecture language models, which matches or exceeds the accuracy of leading full-attention models while significantly improving generation throughput. Jet-Nemotron is developed using Post Neural Architecture Search (PostNAS), a novel neural architecture exploration pipeline that enables efficient model design. Unlike prior approaches, PostNAS begins with a pre-trained full-attention model and freezes its MLP weights, allowing efficient exploration of attention block designs. The pipeline includes four key components: (1) learning optimal full-attention layer placement and elimination, (2) linear attention block selection, (3) designing new attention blocks, and (4) performing hardware-aware hyperparameter search. Our Jet-Nemotron-2B model achieves comparable or superior accuracy to Qwen3, Qwen2.5, Gemma3, and Llama3.2 across a comprehensive suite of benchmarks while delivering up to 53.6× generation throughput speedup and 6.1× prefilling speedup. It also achieves higher accuracy on MMLU and MMLU-Pro than recent advanced MoE full-attention models, such as DeepSeek-V3-Small and Moonlight, despite their larger scale with 15B total and 2.2B activated parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 529e668e-4c9c-4858-8b5e-b8d0e6e29a11Cited by top-tier papers7
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language ModelsYonggan Fu, Xin Dong, Shizhe Diao, Matthijs Van Keirsbilck et al.NeurIPS 2025 · 19 citations
- Distilling to Hybrid Attention Models via KL-Guided Layer SelectionYanhong Li, Songlin Yang, Shawn Tan, Mayank Mishra et al.ICLR 2026 · 17 citations
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin et al.ICLR 2026 · 6 citations
- Composer: A Search Framework for Hybrid Neural Architecture DesignBilge Acun, Prasoon Sinha, Newsha Ardalani, Sangmin Bae et al.ICLR 2026 · 6 citations
- MDN: Parallelizing Stepwise Momentum for Delta Linear AttentionYulong Huang, Xiang Liu, Hongxiang Huang, Xiaopeng LIN et al.ICML 2026 · 1 citation
Builds on46
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts ArchitecturesVenmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour et al.ICML 2026
- FFN Fusion: Rethinking Sequential Computation in Large Language ModelsAkhiad Bercovich, Mohammad Dabbah, Omri Puny, Ido Galil et al.NeurIPS 2025 · 7 citations
- TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and SpeedupFanxu Meng, Pingzhi Tang, Zengwei Yao, Xing Sun et al.NeurIPS 2025 · 5 citations
- Puzzle: Distillation-Based NAS for Inference-Optimized LLMsAkhiad Bercovich, Tomer Ronen, Talor Abramovich, Nir Ailon et al.ICML 2025
- Mining Tensor/Neuron-Level Sparsity to Maximize Mixture-of-Experts Potential in Post-Training and InferenceWeilin Cai, Le Qin, Shwai He, Junwei Cui et al.ICML 2026
