Fast Inference with Kronecker-Sparse Matrices
Antoine Gonon, Léon Zheng, Pascal Carrivain, Quoc-Tung Le
Abstract
Kronecker-sparse (KS) matrices-whose supports are Kronecker products of identity and all-ones blocks-underpin the structure of Butterfly and Monarch matrices and offer the promise of more efficient models. However, existing GPU kernels for KS matrix multiplication suffer from high data movement costs, with up to 50 % of time spent on memory-bound tensor permutations. We propose a fused, output-stationary GPU kernel that eliminates these overheads, reducing global memory traffic threefold. Across 600 KS patterns, our kernel achieves in FP32 a median speedup of ×1.4 and lowers energy consumption by 15 %. A simple heuristic based on KS pattern parameters predicts when our method outperforms existing ones. We release all code 1 at github.com/PascalCarrivain/ksmm, including a PyTorch-compatible KSLinear layer, and demonstrate in FP32 end-to-end latency reductions of up to 22 % in ViT-S/16 and 16 % in GPT-2 medium.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97845f9e-3b12-49c7-ae71-92993fcd35d7Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Monarch: Expressive Structured Matrices for Efficient and Accurate TrainingTri Dao, Beidi Chen, Nimit Sharad Sohoni, Arjun D. Desai et al.ICML 2022 · 125 citations
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network ModelsBeidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang et al.ICLR 2022 · 94 citations
Related papers
- Compute Better Spent: Replacing Dense Layers with Structured MatricesShikai Qiu, Andres Potapczynski, Marc Anton Finzi, Micah Goldblum et al.ICML 2024 · 26 citations
- Monarch Mixer: A Simple Sub-Quadratic GEMM-Based ArchitectureDaniel Y. Fu, Simran Arora, Jessica Grogan, Isys Johnson et al.NeurIPS 2023 · 80 citations
- Accelerating Block Low-Rank Foundation Model Inference on Memory-Constrained GPUsPierre Abillama, Changwoo Lee, Juechu Dong, David T. Blaauw et al.HPDC 2026
- Fast Kronecker Matrix-Matrix Multiplication on GPUsAbhinav Jangda, Mohit YadavPPoPP 2024 · 2 citations
- Searching for Efficient Linear Layers over a Continuous Space of Structured MatricesAndres Potapczynski, Shikai Qiu, Marc Finzi, Christopher Ferri et al.NeurIPS 2024 · 11 citations
