MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication Accelerators
Seunghwan Sung, Sujin Hur, Sungwoo Kim, Dongho Ha, Yunho Oh, Won Woo Ro
Abstract
Modern GPUs commonly employ specialized matrix multiplication units (MXUs) to accelerate matrix multiplication, the core computation of deep learning workloads. However, it is challenging to exploit the MXUs for GPGPU applications whose fundamental algorithms do not rely on matrix multiplication. Furthermore, an additional programming effort is necessary to tailor existing code or algorithms using dedicated APIs or libraries to utilize MXUs. Therefore, MXUs are often underutilized even when GPUs hunger for higher throughput.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e54ff668-52c4-4a62-8097-c86921c8a111Cited by top-tier papers3
- NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free AttentionTianyi Zhang, Jonah Yi, Bowen Yao, Zhaozhuo Xu et al.NeurIPS 2024 · 25 citations
- Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy EfficiencyHansung Kim, Ruohan Richard Yan, Joshua You, Tieliang Vamber Yang et al.ASPLOS 2025 · 1 citation
- Empowering Vector Architectures for ML: The CAMP Architecture for Matrix MultiplicationMohammadreza Esmali Nojehdeh, Hossein Mokhtarnia, Julian Pavon, Narcís Rodas et al.MICRO 2025 · 1 citation
Related papers
- HyTiS: Hybrid Tile Scheduling for GPU GEMM with Enhanced Wave Utilization and Cache LocalityZheng Zhang, Hulin Wang, Hongming Xu, Donglin Yang et al.SC 2025 · 4 citations
- SIMD2: a generalized matrix instruction set for accelerating tensor computation beyond GEMMYunan Zhang, Po-An Tsai, Hung-Wei TsengISCA 2022 · 6 citations
- M3XU: Achieving High-Precision and Complex Matrix Multiplication with Low-Precision MXUsDongho Ha, Yunan Zhang, Chen-Chien Kao, Christopher J. Hughes et al.SC 2024 · 3 citations
- FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor CoresJinliang Shi, Shigang Li, Youxuan Xu, Rongtian Fu et al.PPoPP 2025 · 18 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
