CHOPPER: A Compiler Infrastructure for Programmable Bit-serial SIMD Processing Using Memory in DRAM
Xiangjun Peng, Yaohua Wang, Ming-Chang Yang
Abstract
Increasing interests in Bit-serial SIMD Processing-Using-DRAM (PUD) architectures amplify the needs for a compiler to automate code generation, credited to their ultra-wide SIMD width and reduction of data movements. The state-of-the-art Bit-serial SIMD PUD architectures (1) only provide assembly SIMD programming interfaces, which heavily saddles with programmers to exploit the ultra-wide SIMD width on these architectures; and (2) encapsulate 1-bit operations into multi-bit abstractions, which incurs a granularity mismatch and restricts the optimization space to minimize data movements.We present CHOPPER, a new compiler infrastructure to make Bit-serial SIMD PUD more programmable and efficient. For the better programmability, the design of CHOPPER (1) exploits bit-slicing compilers to enable automatic memory allocation and code generation, from naturally-expressive codes (i.e. similar to Parallel Haskell) into the "SIMD-Within-A-Register"-style codes; and (2) introduces a new abstraction called "Virtual Code Emitter", to make Bit-serial SIMD PUD architecture exploit Memory-Level Parallelism (i.e. Bank or Subarray) more effectively. For the better efficiency, we propose three novel optimizations for CHOPPER to better exploit the potentials of Bit-serial SIMD PUD architectures, which (1) minimize the amount of intra-subarray data movements; and (2) mitigate the overheads of spilling data outside Bit-serial SIMD PUD architectures. These optimizations can greatly improve the overall efficiency of Bit-serial SIMD PUD architectures. We also discuss (1) the limitations of the current CHOPPER; and (2) the potentials of CHOPPER for other types of Processing-In-Memory architectures.We evaluate CHOPPER by hosting it on three state-of-the-art Bit-serial SIMD PUD architectures. We compare CHOPPER-generated codes against the state-of-the-art hands-tuned codes for Bit-serial SIMD PUD architectures. We highlight that, averaged across 16 real-world workloads from 4 PUD-friendly application domains, CHOPPER achieves (A) 1.20X, 1.29X and 1.26X speedup when data can fit within DRAM subarrays; and (B) 12.61X, 9.05X and 9.81X speedup when data need to spill to the secondary storage, on Ambit [50], ELP2IM [56] and SIMDRAM [22], compared with hands-tuned codes using the state-of-the-art methodology [22] for Bit-serial SIMD PUD architectures. These performance benefits also accompany with a great reduction of Lines-of-Codes (LoC) in CHOPPER (i.e. by 4.3X less LoCs for hands-tuning a single subarray, and >103X less for hands-tuning all subarrays in a rank). We also perform breakdown and sensitivity studies of CHOPPER, to better understand its source benefits and examine its robustness under various architectural features.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff9f5304-9c4a-461e-b568-a910f0571606Cited by top-tier papers5
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John et al.ASPLOS 2023 · 20 citations
- BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point CloudsMeng Han, Liang Wang, Limin Xiao, Hao Zhang et al.ISCA 2024 · 14 citations
- CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory ParadigmsAsif Ali Khan, Hamid Farzaneh, Karl Friedrich Alexander Friebel, Clément Fournier et al.ASPLOS 2024 · 7 citations
- Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesRakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta et al.HPCA 2026 · 2 citations
- Count2Multiply: Reliable In-Memory High-Radix CountingJoão Paulo C. de Lima, Benjamin F. Morris III, Asif Ali Khan, Jerónimo Castrillón et al.HPCA 2026
Builds on10
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira et al.ASPLOS 2021 · 182 citations
- ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAMXin Xin, Youtao Zhang, Jun YangHPCA 2020 · 84 citations
- SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory SystemsMaciej Besta, Raghavendra Kanakagiri, Grzegorz Kwasniewski, Rachata Ausavarungnirun et al.MICRO 2021 · 78 citations
- FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and CachingYaohua Wang, Lois Orosa, Xiangjun Peng, Yang Guo et al.MICRO 2020 · 72 citations
- pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup TablesJoão Dinis Ferreira, Gabriel Falcão, Juan Gómez-Luna, Mohammed Alser et al.MICRO 2022 · 60 citations
Related papers
- MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data ComputingGeraldo F. Oliveira, Ataberk Olgun, Abdullah Giray Yaglikçi, F. Nisa Bostanci et al.HPCA 2024 · 44 citations
- UniNDP: A Unified Compilation and Simulation Tool for Near DRAM Processing ArchitecturesTongxin Xie, Zhenhua Zhu, Bing Li, Yukai He et al.HPCA 2025 · 9 citations
- PUMICE: Processing-using-Memory Integration with a Scalar Pipeline for Symbiotic ExecutionSocrates S. Wong, Cecilio C. Tamarit, José F. MartínezDAC 2023 · 5 citations
- ATiM: Autotuning Tensor Programs for Processing-in-DRAMYongwon Shin, Dookyung Kang, Hyojin SungISCA 2025 · 2 citations
- Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU CoresZhongcheng Zhang, Yan Ou, Ying Liu, Chenxi Wang et al.ASPLOS 2023 · 4 citations
