Debunking the CUDA Myth Towards GPU-based AI Systems: Evaluation of the Performance and Programmability of Intel's Gaudi NPU for AI Model Serving
Yunjae Lee, Juntaek Lim, Jehyeon Bang, Eunyeong Cho, Huijong Jeong, Taesu Kim, Hyungjun Kim, Joonhyung Lee, Jinseop Im, Ranggi Hwang, Se Jung Kwon, Dongsoo Lee, Minsoo Rhu
摘要
This paper presents a comprehensive evaluation of Intel Gaudi NPUs as an alternative to NVIDIA GPUs, which is currently the de facto standard in AI system design. First, we create a suite of microbenchmarks to compare Intel Gaudi-2 with NVIDIA A100, showing that Gaudi-2 achieves competitive performance not only in primitive AI compute, memory, and communication operations but also in executing several important AI workloads end-to-end. We then assess Gaudi NPU's programmability by discussing several software-level optimization strategies to employ for implementing critical FBGEMM operators and vLLM, evaluating their efficiency against GPU-optimized counterparts. Results indicate that Gaudi-2 achieves energy efficiency comparable to A100, though there are notable areas for improvement in terms of software maturity. Overall, we conclude that, with effective integration into highlevel AI frameworks, Gaudi NPUs could challenge NVIDIA GPU's dominance in the AI server market, though further improvements are necessary to fully compete with NVIDIA's robust software ecosystem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain 等WWW 2021 · 被引用 793 次
相关 Paper
- Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power SensorZeyu Yang, Karel Adámek, Wesley ArmourSC 2024 · 被引用 38 次
- Characterizing Performance, Power, and Energy of AMD CDNA3 GPU FamilyBagus Hanindhito, Bhavesh PatelSC 2025 · 被引用 2 次
- Tandem Processor: Grappling with Emerging Operators in Neural NetworksSoroush Ghodrati, Sean Kinzer, Hanyang Xu, Rohan Mahapatra 等ASPLOS 2024 · 被引用 21 次
- XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM WorkloadsMingcong Song, Xinru Tang, Fengfan Hou, Jing Li 等ASPLOS 2026
- Guardain: Protecting Emerging Generative AI Workloads on Heterogeneous NPUAritra Dhar, Clément Thorens, Lara Magdalena Lazier, Lukas CavigelliS&P 2025
