SC2021Top-tier venue
LIBSHALOM: optimizing small and irregular-shaped matrix multiplications on ARMv8 multi-cores
Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su, Zheng Wang
Abstract
General Matrix Multiplication (GEMM) is a key subroutine in highperformance computing. While the mainstream linear algebra libraries can deliver high performance on large and regular-shaped GEMM, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in new HPC applications. Some of the recent works in this direction have made promising progress on x86 architectures and GPUs but still leave much room for improvement on emerging HPC hardware built upon the ARMv8 architecture. We present LibShalom, an open-source library for optimizing small and irregular-shaped GEMMs, explicitly targeting the ARMv8 architecture. LibShalom builds upon the classical Goto algorithm but tailors it to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluate LibShalom by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show that LibShalom can consistently outperform existing solutions across GEMM workloads and hardware architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm ArchitecturesDu Wu, Jintao Meng, Wenxi Zhu, Minwen Deng et al.SC 2024 · 14 citations
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
- AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 ProcessorsHaoyuan Gui, Xiaoyu Zhang, Yifan Zhang, Ximeng Fu et al.AAAI 2026
Builds on2
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- Co-design for A64FX manycore processor and "Fugaku"Mitsuhisa Sato, Yutaka Ishikawa, Hirofumi Tomita, Yuetsu Kodama et al.SC 2020 · 112 citations
Related papers
- ASM-SpMM: Unleashing the Potential of Arm SME for Sparse Matrix Multiplication AccelerationJiazhi Jiang, Xijia Yao, Jiayu Chen, Jinhui Wei et al.PPoPP 2026
- HSMU-SpGEMM: Achieving High Shared Memory Utilization for Parallel Sparse General Matrix-Matrix Multiplication on Modern GPUsMin Wu, Huizhang Luo, Fenfang Li, Yiran Zhang et al.HPCA 2025 · 3 citations
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari et al.HPCA 2021 · 66 citations
- A Hardware-Software Design Framework for SpMV Acceleration with Flexible Access Pattern PortfolioZhenyu Wu, Maolin Wang, Hayden Kwok-Hay SoHPCA 2025 · 1 citation
