Puzzle: Distillation-Based NAS for Inference-Optimized LLMs
Akhiad Bercovich, Tomer Ronen, Talor Abramovich, Nir Ailon, Nave Assaf, Mohammad Dabbah, Ido Galil, Amnon Geifman, Yonatan Geifman, Izhak Golan, Netanel Haber, Ehud Karpas
摘要
Large language models (LLMs) offer remarkable capabilities, yet their high inference costs restrict wider adoption. While increasing parameter counts improves accuracy, it also broadens the gap between state-of-the-art capabilities and practical deployability. We present Puzzle, a hardware-aware framework that accelerates the inference of LLMs while preserving their capabilities. Using neural architecture search (NAS) at a large-scale, Puzzle optimizes models with tens of billions of parameters. Our approach utilizes blockwise local knowledge distillation (BLD) for parallel architecture exploration and employs mixed-integer programming for precise constraint optimization. We showcase our framework's impact via Llama-3.1-Nemotron-51B-Instruct (Nemotron-51B) and Llama-3.3-Nemotron-49B, two publicly available models derived from Llama-70B-Instruct. Both models achieve a 2.17× inference throughput speedup, fitting on a single NVIDIA H100 GPU while retaining 98.4% of the original model's benchmark accuracies. These are the most accurate models supporting single H100 GPU inference with large batch sizes, despite training on 45B tokens at most, far fewer than the 15T used to train Llama-70B. Lastly, we show that lightweight alignment on these derived models allows them to surpass the parent model in specific capabilities. Our work establishes that powerful LLM models can be optimized for efficient deployment with only negligible loss in quality, underscoring that inference performance, not parameter count alone, should guide model selection. LLMs require a substantial amount of parameters for their training process to converge easily and achieve better generalization [29, 25, 5, 8] . This overparameterization not only facilitates optimization, but also provides greater capacity to store knowledge and learn complex patterns across diverse tasks, explaining why larger models consistently demonstrate superior performance [29, 25] . However, once trained, many parameters and computations turn out to be redundant for inference, as evidenced by the success of various computational efficiency techniques [21, 9, 58, 7, 40, 27, 3 ]. Yet, LLM architectures remain largely uniform, comprising repeated identical layers, with little consideration given to balancing each block's computational cost against its contribution to overall model predictive performance-a *These authors contributed equally. Other co-authors are listed alphabetically. design choice primarily driven by training stability and ease of scaling rather than inference efficiency. This work addresses how to transform a trained LLM from a structure suited for training into one optimized for efficient inference on specific hardware (such as H100), while preserving its accumulated knowledge and predictive performance. Given a "parent model", our approach explores a large search space of architecture configurations to identify efficient options tailored to meet specific hardware and task-related constraints. This exploration requires a method to reliably estimate the performance of each potential configuration, allowing us to identify models that balance efficiency and accuracy for deployment. MHA FFN MHA Linear no-op FFN MQA Linear GQA FFN GQA Linear GQA Linear MQA FFN Step 1: Crafting the "puzzle pieces" Applying block-wise local distillation to every alternative subblock replacement in parallel and scoring its quality and inference cost to build a "library" of blocks. Step 2: Assembling the puzzle architecture Utilizing Mixed-Integer-Programming to assemble a heterogeneous architecture that optimizes quality under constraints such as throughput, latency and memory usage. Step 3: Uptraining The reassembled architecture is trained with global Knowledge-Distillation to strengthen interblock compatibility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture SearchYuxian Gu, Qinghao Hu, Haocheng Xi, Junyu Chen 等NeurIPS 2025 · 被引用 39 次
- Hankel Singular Value Regularization for Highly Compressible State Space ModelsPaul Schwerdtner, Jules Berman, Benjamin PeherstorferNeurIPS 2025 · 被引用 3 次
- Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative DecodingJinze Li, Yixing Xu, Haiduo Huang, Xuanwu Yin 等ICML 2025
- WAVE: Window-Aware Vocabulary-Efficient Early-Exit for Training-Free LLM AccelerationSeonggeun Kim, Gilha lee, Hyun KimICML 2026
它引用的顶会 Paper28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context SwitchKinman Lei, Yuyang Jin, Mingshu Zhai, Kezhao Huang 等USENIX ATC 2024 · 被引用 19 次
- Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMsSong Bian, Tao Yu, Shivaram Venkataraman, Youngsuk ParkICLR 2026 · 被引用 3 次
- Learning Semi-Structured Sparsity for LLMs via Shared and Context-Aware HypernetworkLu Sun, Jun SakumaICLR 2026
- FFN Fusion: Rethinking Sequential Computation in Large Language ModelsAkhiad Bercovich, Mohammad Dabbah, Omri Puny, Ido Galil 等NeurIPS 2025 · 被引用 7 次
- NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMMConnor Holmes, Minjia Zhang, Yuxiong He, Bo WuNeurIPS 2021 · 被引用 29 次
