VeGen: a vectorizer generator for SIMD and beyond
Yishen Chen, Charith Mendis, Michael Carbin, Saman P. Amarasinghe
摘要
Vector instructions are ubiquitous in modern processors. Traditional compiler auto-vectorization techniques have focused on targeting single instruction multiple data (SIMD) instructions. However, these auto-vectorization techniques are not sufficiently powerful to model non-SIMD vector instructions, which can accelerate applications in domains such as image processing, digital signal processing, and machine learning. To target non-SIMD instruction, compiler developers have resorted to complicated, ad hoc peephole optimizations, expending significant development time while still coming up short. As vector instruction sets continue to rapidly evolve, compilers cannot keep up with these new hardware capabilities.
In this paper, we introduce Lane Level Parallelism (LLP), which captures the model of parallelism implemented by both SIMD and non-SIMD vector instructions. We present VeGen, a vectorizer generator that automatically generates a vectorization pass to uncover target-architecture-specific LLP in programs while using only instruction semantics as input. VeGen decouples, yet coordinates automatically generated target-specific vectorization utilities with its target-independent vectorization algorithm. This design enables us to systematically target non-SIMD vector instructions that until now require ad hoc coordination between different compiler stages. We show that VeGen can use non-SIMD vector instructions effectively, for example, getting speedup 3× (compared to LLVM's vectorizer) on x265's idct4 kernel.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hardware abstractionSize Zheng, Renze Chen, Anjiang Wei, Yicheng Jin 等ISCA 2022 · 被引用 63 次
- A Tensor Compiler with Automatic Data Packing for Simple and Efficient Fully Homomorphic EncryptionAleksandar Krastev, Nikola Samardzic, Simon Langowski, Srinivas Devadas 等PLDI 2024 · 被引用 23 次
- Domain specific run time optimization for software data planesSebastiano Miano, Alireza Sanaee, Fulvio Risso, Gábor Rétvári 等ASPLOS 2022 · 被引用 20 次
- Automatic Generation of Vectorizing Compilers for Customizable Digital Signal ProcessorsSamuel Thomas, James BornholtASPLOS 2024 · 被引用 16 次
- Vector instruction selection for digital signal processors using program synthesisMaaz Bin Safeer Ahmad, Alexander J. Root, Andrew Adams, Shoaib Kamil 等ASPLOS 2022 · 被引用 13 次
相关 Paper
- All you need is superword-level parallelism: systematic control-flow vectorization with SLPYishen Chen, Charith Mendis, Saman P. AmarasinghePLDI 2022 · 被引用 20 次
- Boost Linear Algebra Computation Performance via Efficient VNNI UtilizationHao Zhou, Qiukun Han, Heng Shi, Yalin Zhang 等ASPLOS 2024 · 被引用 1 次
- Hydride: A Retargetable and Extensible Synthesis-based Compiler for Modern Hardware ArchitecturesAkash Kothari, Abdul Rafae Noor, Muchen Xu, Hassam Uddin 等ASPLOS 2024 · 被引用 8 次
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 被引用 27 次
- Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline ParallelismQuan M. Nguyen, Daniel SánchezHPCA 2023 · 被引用 7 次
