Boost Linear Algebra Computation Performance via Efficient VNNI Utilization
Hao Zhou, Qiukun Han, Heng Shi, Yalin Zhang, Jianguo Yao
Abstract
Intel's Vector Neural Network Instruction (VNNI) provides higher efficiency on calculating dense linear algebra (DLA) computations than conventional SIMD instructions. However, existing auto-vectorizers frequently deliver suboptimal utilization of VNNI by either failing to recognize VNNI's unique computation pattern at the innermost loops/basic blocks, or producing inferior code through constrained and rudimentary peephole optimizations/pattern matching techniques. Auto-tuning frameworks might generate proficient code but are hampered by the necessity for sophisticated pattern templates and extensive search processes.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPUJiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang et al.DAC 2020 · 15 citations
- VeGen: a vectorizer generator for SIMD and beyondYishen Chen, Charith Mendis, Michael Carbin, Saman P. AmarasingheASPLOS 2021 · 43 citations
- VIA: A Smart Scratchpad for Vector Units with Application to Sparse Matrix ComputationsJulian Pavon, Iván Vargas Valdivieso, Adrián Barredo, Joan Marimon et al.HPCA 2021 · 21 citations
- BiSon-e: a lightweight and high-performance accelerator for narrow integer linear algebra computing on the edgeEnrico Reggiani, Cristóbal Ramírez Lazo, Roger Figueras Bagué, Adrián Cristal et al.ASPLOS 2022 · 11 citations
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev et al.ASPLOS 2021 · 52 citations
