Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
Yuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen, Shang Yang, Song Han, Han Cai
摘要
We present Jet-Nemotron, a new family of hybrid-architecture language models, which matches or exceeds the accuracy of leading full-attention models while significantly improving generation throughput. Jet-Nemotron is developed using Post Neural Architecture Search (PostNAS), a novel neural architecture exploration pipeline that enables efficient model design. Unlike prior approaches, PostNAS begins with a pre-trained full-attention model and freezes its MLP weights, allowing efficient exploration of attention block designs. The pipeline includes four key components: (1) learning optimal full-attention layer placement and elimination, (2) linear attention block selection, (3) designing new attention blocks, and (4) performing hardware-aware hyperparameter search. Our Jet-Nemotron-2B model achieves comparable or superior accuracy to Qwen3, Qwen2.5, Gemma3, and Llama3.2 across a comprehensive suite of benchmarks while delivering up to 53.6× generation throughput speedup and 6.1× prefilling speedup. It also achieves higher accuracy on MMLU and MMLU-Pro than recent advanced MoE full-attention models, such as DeepSeek-V3-Small and Moonlight, despite their larger scale with 15B total and 2.2B activated parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language ModelsYonggan Fu, Xin Dong, Shizhe Diao, Matthijs Van Keirsbilck 等NeurIPS 2025 · 被引用 19 次
- Distilling to Hybrid Attention Models via KL-Guided Layer SelectionYanhong Li, Songlin Yang, Shawn Tan, Mayank Mishra 等ICLR 2026 · 被引用 17 次
- UniQL: Unified Quantization and Low-rank Compression for Adaptive Edge LLMsHung-Yueh Chiang, Chi-Chih Chang, Yu-Chen Lu, Chien-Yu Lin 等ICLR 2026 · 被引用 6 次
- Composer: A Search Framework for Hybrid Neural Architecture DesignBilge Acun, Prasoon Sinha, Newsha Ardalani, Sangmin Bae 等ICLR 2026 · 被引用 6 次
- MDN: Parallelizing Stepwise Momentum for Delta Linear AttentionYulong Huang, Xiang Liu, Hongxiang Huang, Xiaopeng LIN 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper46
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Revisiting Efficiency–Accuracy Scaling in Mixture-of-Experts ArchitecturesVenmugil Elango, Nidhi Bhatia, Roger Waleffe, Rasoul Shafipour 等ICML 2026
- FFN Fusion: Rethinking Sequential Computation in Large Language ModelsAkhiad Bercovich, Mohammad Dabbah, Omri Puny, Ido Galil 等NeurIPS 2025 · 被引用 7 次
- TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and SpeedupFanxu Meng, Pingzhi Tang, Zengwei Yao, Xing Sun 等NeurIPS 2025 · 被引用 5 次
- Puzzle: Distillation-Based NAS for Inference-Optimized LLMsAkhiad Bercovich, Tomer Ronen, Talor Abramovich, Nir Ailon 等ICML 2025
- Mining Tensor/Neuron-Level Sparsity to Maximize Mixture-of-Experts Potential in Post-Training and InferenceWeilin Cai, Le Qin, Shwai He, Junwei Cui 等ICML 2026
