Interleaved Multi-Vectorizing
Zhuhe Fang, Beilei Zheng, Chuliang Weng
摘要
SIMD is an instruction set in mainstream processors, which provides the data level parallelism to accelerate the performance of applications. However, its advantages diminish when applications suffer from heavy cache misses. To eliminate cache misses in SIMD vectorization, we present interleaved multi-vectorizing (IMV) in this paper. It interleaves multiple execution instances of vectorized code to hide memory access latency with more computation. We also propose residual vectorized states to solve the control flow divergence in vectorization. IMV can make full use of the data parallelism in SIMD and the memory level parallelism through prefetching. It reduces cache misses, branch misses and computation overhead to significantly speed up the performance of pointer-chasing applications, and it can be applied to executing entire query pipelines. As experimental results show, IMV achieves up to 4.23X and 3.17X better performance compared with the pure scalar implementation and the pure SIMD vectorization, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CoroBase: Coroutine-Oriented Main-Memory Database EngineYongjun He, Jiacheng Lu, Tianzheng WangVLDB 2021 · 被引用 42 次
- The Case for Learned In-Memory JoinsIbrahim Sabek, Tim KraskaVLDB 2023 · 被引用 27 次
- The Art of Latency Hiding in Modern Database EnginesKaisong Huang, Tianzheng Wang, Qingqing Zhou, Qingzhong MengVLDB 2024 · 被引用 23 次
- Charting the Design Space of Query Execution using VOILATim Gubner, Peter BonczVLDB 2021 · 被引用 17 次
相关 Paper
- Vector RunaheadAjeya Naithani, Sam Ainsworth, Timothy M. Jones, Lieven EeckhoutISCA 2021 · 被引用 27 次
- Decoupled Vector RunaheadAjeya Naithani, Jaime Roelandts, Sam Ainsworth, Timothy M. Jones 等MICRO 2023 · 被引用 15 次
- Multi-Dimensional Vector ISA Extension for Mobile In-Cache ComputingAlireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu 等HPCA 2025 · 被引用 4 次
- Database Processing-in-Memory: An Experimental StudyTiago Rodrigo Kepe, Eduardo C. de Almeida, Marco A. Z. AlvesVLDB 2020 · 被引用 22 次
- PUMICE: Processing-using-Memory Integration with a Scalar Pipeline for Symbiotic ExecutionSocrates S. Wong, Cecilio C. Tamarit, José F. MartínezDAC 2023 · 被引用 5 次
