SC2024Top-tier venue
Unlocking High Performance with Low-Bit NPUs and CPUs for Highly Optimized HPL-MxP on Cloud Brain II
Weicheng Xue, Kai Yang, Yongxiang Liu, Dengdong Fan, Pengxiang Xu, Yonghong Tian
Abstract
Mix-precision computation is crucial for artificial intelligence and scientific computing applications. However, as novel chips with innovative architectures emerge, harnessing their computational capabilities presents significant challenges. While existing algorithms for the HPL-MxP LU factorization excel on homogeneous systems, they often encounter difficulties on specialized heterogeneous architectures. This deficiency arises from inadequate optimization for computation, memory access, and communication, hindering effective mixed-precision acceleration. This work introduces an algorithm-hardware co-optimization approach for LU factorization on specialized NPUs and CPUs, leveraging their unique architectures. A novel multi-iteration fusion method for general matrix multiplication is proposed, strategically designed to maximize on-chip L1 buffer utilization, effectively overcoming the notorious “memory wall”. Additionally, a multi-stage, multi-level heterogeneous pipeline for LU factorization in an accelerator-CPU cloud environment is presented, where compute-intensive matrix multiplications are offloaded to NPUs while CPUs handle the remaining tasks. The co-optimization approach fosters deep collaboration between CPUs and accelerators, thereby unlocking enhanced performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- 5 ExaFlop/s HPL-MxP Benchmark with Linear Scalability on the 40-Million-Core Sunway SupercomputerRongfen Lin, Xinhui Yuan, Wei Xue, Wanwang Yin et al.SC 2023 · 11 citations
- SpaHet: A Software/Hardware Co-design for Accelerating Heterogeneous-Sparsity based Sparse Matrix MultiplicationHaoqin Huang, Pengcheng Yao, Zhaozeng An, Yufei Sun et al.DAC 2024 · 3 citations
- Spatula: A Hardware Accelerator for Sparse Matrix FactorizationAxel Feldmann, Daniel SánchezMICRO 2023 · 9 citations
- MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsSeunghwan Sung, Sujin Hur, Sungwoo Kim, Dongho Ha et al.MICRO 2023 · 5 citations
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
