Boost Linear Algebra Computation Performance via Efficient VNNI Utilization
Hao Zhou, Qiukun Han, Heng Shi, Yalin Zhang, Jianguo Yao
2024年份
1被引次数
摘要
Intel's Vector Neural Network Instruction (VNNI) provides higher efficiency on calculating dense linear algebra (DLA) computations than conventional SIMD instructions. However, existing auto-vectorizers frequently deliver suboptimal utilization of VNNI by either failing to recognize VNNI's unique computation pattern at the innermost loops/basic blocks, or producing inferior code through constrained and rudimentary peephole optimizations/pattern matching techniques. Auto-tuning frameworks might generate proficient code but are hampered by the necessity for sophisticated pattern templates and extensive search processes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- A History-Based Auto-Tuning Framework for Fast and High-Performance DNN Design on GPUJiandong Mu, Mengdi Wang, Lanbo Li, Jun Yang 等DAC 2020 · 被引用 15 次
- VeGen: a vectorizer generator for SIMD and beyondYishen Chen, Charith Mendis, Michael Carbin, Saman P. AmarasingheASPLOS 2021 · 被引用 43 次
- VIA: A Smart Scratchpad for Vector Units with Application to Sparse Matrix ComputationsJulian Pavon, Iván Vargas Valdivieso, Adrián Barredo, Joan Marimon 等HPCA 2021 · 被引用 21 次
- BiSon-e: a lightweight and high-performance accelerator for narrow integer linear algebra computing on the edgeEnrico Reggiani, Cristóbal Ramírez Lazo, Roger Figueras Bagué, Adrián Cristal 等ASPLOS 2022 · 被引用 11 次
- Analytical characterization and design space exploration for optimization of CNNsRui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev 等ASPLOS 2021 · 被引用 52 次
