Squeezing Operator Performance Potential for the Ascend Architecture
Yuhang Zhou, Zhibin Wang, Guyue Liu, Shipeng Li, Xi Lin, Zibo Wang, Yongzhong Wang, Fuchun Wei, Jingyi Zhang, Zhiheng Hu, Yanlin Liu, Chunsheng Li
摘要
With the rise of deep learning, many companies have developed domain-specific architectures (DSAs) optimized for AI workloads, with Ascend being a representative. To fully realize the operator performance on Ascend, effective analysis and optimization is urgently needed. Compared to GPU, Ascend requires users to manage operations manually, leading to complex performance issues that require precise analysis. However, existing roofline models face challenges of visualization complexity and inaccurate performance assessment. To address these needs, we introduce a component-based roofline model that abstracts components to capture operator performance, thereby effectively identifying bottleneck components. Furthermore, through practical operator optimization case studies, we illustrate a comprehensive process of optimization based on roofline analysis, summarizing common performance issues and optimization strategies. Finally, extensive end-to-end optimization experiments demonstrate significant model speed improvements, ranging from 1.07× to 2.15×, along with valuable insights from practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CoDec: Prefix-Shared Decoding Kernel for LLMsZhibin Wang, Rui Ning, Chao Fang, Zhonghui Zhang 等SIGMOD 2026 · 被引用 8 次
- Accelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and OptimizationYuhang Zhou, Zibo Wang, Zhibin Wang, Ruyi Zhang 等USENIX ATC 2025 · 被引用 6 次
它引用的顶会 Paper6
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal 等PLDI 2021 · 被引用 166 次
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama 等ICLR 2020 · 被引用 159 次
- Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network TrainingGeoffrey X. Yu, Yubo Gao, Pavel Golikov, Gennady PekhimenkoUSENIX ATC 2021 · 被引用 108 次
- AKG: automatic kernel generation for neural processing units using polyhedral transformationsJie Zhao, Bojie Li, Wang Nie, Zhen Geng 等PLDI 2021 · 被引用 81 次
相关 Paper
- The Configuration Wall: Characterization and Elimination of Accelerator Configuration OverheadJosse Van Delm, Anton Lydike, Joren Dumoulin, Jonas Crols 等ASPLOS 2026
- NASGuard: A Novel Accelerator Architecture for Robust Neural Architecture Search (NAS) NetworksXingbin Wang, Boyan Zhao, Rui Hou, Amro Awad 等ISCA 2021 · 被引用 9 次
- Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyZibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang 等ASPLOS 2025 · 被引用 9 次
- AIO: An Abstraction for Performance Analysis Across Diverse Accelerator ArchitecturesJoseph Rogers, Taha Soliman, Magnus JahreISCA 2024 · 被引用 5 次
- XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM WorkloadsMingcong Song, Xinru Tang, Fengfan Hou, Jing Li 等ASPLOS 2026
