Squeezing Operator Performance Potential for the Ascend Architecture
Yuhang Zhou, Zhibin Wang, Guyue Liu, Shipeng Li, Xi Lin, Zibo Wang, Yongzhong Wang, Fuchun Wei, Jingyi Zhang, Zhiheng Hu, Yanlin Liu, Chunsheng Li
Abstract
With the rise of deep learning, many companies have developed domain-specific architectures (DSAs) optimized for AI workloads, with Ascend being a representative. To fully realize the operator performance on Ascend, effective analysis and optimization is urgently needed. Compared to GPU, Ascend requires users to manage operations manually, leading to complex performance issues that require precise analysis. However, existing roofline models face challenges of visualization complexity and inaccurate performance assessment. To address these needs, we introduce a component-based roofline model that abstracts components to capture operator performance, thereby effectively identifying bottleneck components. Furthermore, through practical operator optimization case studies, we illustrate a comprehensive process of optimization based on roofline analysis, summarizing common performance issues and optimization strategies. Finally, extensive end-to-end optimization experiments demonstrate significant model speed improvements, ranging from 1.07× to 2.15×, along with valuable insights from practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01cf2add-a650-4878-9ea0-bfc1a78a828bCited by top-tier papers2
- CoDec: Prefix-Shared Decoding Kernel for LLMsZhibin Wang, Rui Ning, Chao Fang, Zhonghui Zhang et al.SIGMOD 2026 · 8 citations
- Accelerating Model Training on Ascend Chips: An Industrial System for Profiling, Analysis and OptimizationYuhang Zhou, Zibo Wang, Zhibin Wang, Ruyi Zhang et al.USENIX ATC 2025 · 6 citations
Builds on6
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- DNNFusion: accelerating deep neural networks execution with advanced operator fusionWei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal et al.PLDI 2021 · 166 citations
- Mixed Precision DNNs: All you need is a good parametrizationStefan Uhlich, Lukas Mauch, Fabien Cardinaux, Kazuki Yoshiyama et al.ICLR 2020 · 159 citations
- Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network TrainingGeoffrey X. Yu, Yubo Gao, Pavel Golikov, Gennady PekhimenkoUSENIX ATC 2021 · 108 citations
- AKG: automatic kernel generation for neural processing units using polyhedral transformationsJie Zhao, Bojie Li, Wang Nie, Zhen Geng et al.PLDI 2021 · 81 citations
Related papers
- The Configuration Wall: Characterization and Elimination of Accelerator Configuration OverheadJosse Van Delm, Anton Lydike, Joren Dumoulin, Jonas Crols et al.ASPLOS 2026
- NASGuard: A Novel Accelerator Architecture for Robust Neural Architecture Search (NAS) NetworksXingbin Wang, Boyan Zhao, Rui Hou, Amro Awad et al.ISCA 2021 · 9 citations
- Using Analytical Performance/Power Model and Fine-Grained DVFS to Enhance AI Accelerator Energy EfficiencyZibo Wang, Yijia Zhang, Fuchun Wei, Bingqiang Wang et al.ASPLOS 2025 · 9 citations
- AIO: An Abstraction for Performance Analysis Across Diverse Accelerator ArchitecturesJoseph Rogers, Taha Soliman, Magnus JahreISCA 2024 · 5 citations
- XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM WorkloadsMingcong Song, Xinru Tang, Fengfan Hou, Jing Li et al.ASPLOS 2026
