SC2025Top-tier venue
HyTiS: Hybrid Tile Scheduling for GPU GEMM with Enhanced Wave Utilization and Cache Locality
Zheng Zhang, Hulin Wang, Hongming Xu, Donglin Yang, Xiaobo Zhou, Dazhao Cheng
Abstract
General matrix-matrix multiplication (GEMM) is a fundamental operation in both deep learning and scientific computing. To accelerate these workloads, GPUs with a large number of streaming multiprocessors (SMs) are widely used. However, as modern GPUs scale in core count and adopt larger tile sizes, the wave quantization problem induced by partially filled waves results in growing hardware underutilization and substantially degraded performance. Existing solutions to this problem often suffer from low execution efficiency or introduce additional synchronization overhead.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 73b431f7-9bf0-4235-a4a8-1e86745d5f5dCited by top-tier papers1
Ask how each one uses itRelated papers
- MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsSeunghwan Sung, Sujin Hur, Sungwoo Kim, Dongho Ha et al.MICRO 2023 · 5 citations
- TileSpGEMM: a tiled algorithm for parallel sparse general matrix-matrix multiplication on GPUsYuyao Niu, Zhengyang Lu, Haonan Ji, Shuhui Song et al.PPoPP 2022 · 66 citations
- KAMI: Communication-Avoiding General Matrix Multiplication within a Single GPUHemeng Wang, Yang Du, Sidu Li, Xiaowen Tian et al.SC 2025 · 4 citations
- HSMU-SpGEMM: Achieving High Shared Memory Utilization for Parallel Sparse General Matrix-Matrix Multiplication on Modern GPUsMin Wu, Huizhang Luo, Fenfang Li, Yiran Zhang et al.HPCA 2025 · 3 citations
- HARP: Hardware-Based Pseudo-Tiling for Sparse Matrix Multiplication AcceleratorJinkwon Kim, Myeongjae Jang, Haejin Nam, Soontae KimMICRO 2023 · 12 citations
