High Accuracy Matrix Computations on Neural Engines: A Study of QR Factorization and its Applications
Shaoshuai Zhang, Elaheh Baharlouei, Panruo Wu
Abstract
Fueled by the surge of ever expanding successful applications of deep neural networks and the great computational power demanded, modern computer processors and accelerators are beginning to offer half precision floating point arithmetic support, and special units (neural engines) such as NVIDIA TensorCore on GPU and Google Tensor Processing Unit (TPU) to accelerate the training and prediction of deep neural networks. It remains unclear how neural engines can be profitably used in application other than neural networks. In this paper we present an endeavor of accelerating and stabilizing a fundamental matrix factorization on neural enginesthe QR factorization-which may open doors to much wider relevance to scientific, engineering, and data sciences. We show that traditional Householder QR algorithms and implementations do not have the necessary data locality, parallelism, accuracy and robustness on neural engines which are characterized by extreme speed and low precision/range.
We demonstrate that neural engines can be effectively used to accelerate matrix computations (QR 3.0x-14.6x speedup compared to cuSOLVER, reaching up to 36.6TFLOPS); however different algorithms (recursive Gram-Schmidt) are needed to expose more locality and parallelism, even at the cost of increased computations. Moreover, scaling, iterative refinement, and other safeguarding procedures are also needed to regain accuracy and avoid overflowing. Our experience seems to suggest that presently with neural engines the matrix factorizations (QR, LU, Cholesky) are best to be co-designed with its applications (linear solver, least square, orthogonalization, SVD etc) to achieve high performance and adequate accuracy and reliability, rather than used as a black-box.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fccde857-03ae-4f02-be7c-34a29b00aa8eCited by top-tier papers3
- AmgT: Algebraic Multigrid Solver on Tensor CoresYuechen Lu, Lijie Zeng, Tengcheng Wang, Xu Fu et al.SC 2024 · 17 citations
- HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAHan Huang, Jiabin Xie, Guangnan Feng, Xianwei Zhang et al.SC 2025 · 5 citations
- Spectral Basis Learning for Expressive Graph Neural Networks in Link PredictionNiloofar Azizi, Nils M. Kriege, Nicholas J. A. Harvey, Horst BischofAAAI 2026
Related papers
- TCUDB: Accelerating Database with Tensor ProcessorsYu-Ching Hu, Yuliang Li, Hung-Wei TsengSIGMOD 2022 · 32 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Cooperative Warp Execution in Tensor Core for RISC-V GPGPUAbubakr Nada, Giuseppe Maria Sarda, Erwan LenormandHPCA 2025 · 3 citations
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang et al.DAC 2020 · 12 citations
- Efficient Quantized Sparse Matrix Operations on Tensor CoresShigang Li, Kazuki Osawa, Torsten HoeflerSC 2022 · 27 citations
