Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU Cores
Zhongcheng Zhang, Yan Ou, Ying Liu, Chenxi Wang, Yongbin Zhou, Xiaoyu Wang, Yuyang Zhang, Yucheng Ouyang, Jiahao Shan, Ying Wang, Jingling Xue, Huimin Cui, Xiaobing Feng
Abstract
SIMD extensions are widely adopted in multi-core processors to exploit data-level parallelism. However, when co-running workloads on different cores, compute-intensive workloads cannot take advantage of the underutilized SIMD lanes allocated to memoryintensive workloads, reducing the overall performance. This paper proposes Occamy, a SIMD co-processor that can be shared by multiple CPU cores, so that their co-running workloads can spatially share its SIMD lanes. The key idea is to enable elastic spatial sharing by dynamically partitioning all the SIMD lanes across different workloads based on their phase behaviors, so that each workload may execute in variable-length SIMD mode. We also introduce an Occamy compiler to support such variable-length vectorization by analyzing such phase behaviors and generating the vectorized code that works with varying vector lengths. We demonstrate that Occamy can improve SIMD utilization, and consequently, performance over three representative SIMD architectures, with negligible chip area cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Vectorization for digital signal processors via equality saturationAlexa VanHattum, Rachit Nigam, Vincent T. Lee, James Bornholt et al.ASPLOS 2021 · 57 citations
- VeGen: a vectorizer generator for SIMD and beyondYishen Chen, Charith Mendis, Michael Carbin, Saman P. AmarasingheASPLOS 2021 · 43 citations
- Unlimited Vector Extension with Data Streaming SupportJoao Mario Domingos, Nuno Neves, Nuno Roma, Pedro TomásISCA 2021 · 31 citations
- All you need is superword-level parallelism: systematic control-flow vectorization with SLPYishen Chen, Charith Mendis, Saman P. AmarasinghePLDI 2022 · 20 citations
- REDUCT: Keep it Close, Keep it Cool! : Efficient Scaling of DNN Inference on Multi-core CPUs with Near-Cache ComputeAnant V. Nori, Rahul Bera, Shankar Balachandran, Joydeep Rakshit et al.ISCA 2021 · 17 citations
Related papers
- Occamy: Memory-efficient GPU Compiler for DNN InferenceJaeho Lee, Shinnung Jeong, Seungbin Song, Kunwoo Kim et al.DAC 2023 · 3 citations
- Interleaved Multi-VectorizingZhuhe Fang, Beilei Zheng, Chuliang WengVLDB 2020 · 19 citations
- CHOPPER: A Compiler Infrastructure for Programmable Bit-serial SIMD Processing Using Memory in DRAMXiangjun Peng, Yaohua Wang, Ming-Chang YangHPCA 2023 · 15 citations
- PhaseWeave: Phase-Aware Execution on Heterogeneous Chiplet Architectures for DatacentersJoshua Kim, Chaojie Zhang, Íñigo Goiri, Christopher J. Rossbach et al.ISCA 2026 · 1 citation
- Enhancing Thread-Level Parallelism in Asymmetric Multicores using Transparent Instruction OffloadingJeckson Dellagostin Souza, Madhavan Manivannan, Miquel Pericàs, Antonio Carlos Schneider BeckDAC 2020 · 1 citation
