CHOPPER: A Compiler Infrastructure for Programmable Bit-serial SIMD Processing Using Memory in DRAM
Xiangjun Peng, Yaohua Wang, Ming-Chang Yang
摘要
Increasing interests in Bit-serial SIMD Processing-Using-DRAM (PUD) architectures amplify the needs for a compiler to automate code generation, credited to their ultra-wide SIMD width and reduction of data movements. The state-of-the-art Bit-serial SIMD PUD architectures (1) only provide assembly SIMD programming interfaces, which heavily saddles with programmers to exploit the ultra-wide SIMD width on these architectures; and (2) encapsulate 1-bit operations into multi-bit abstractions, which incurs a granularity mismatch and restricts the optimization space to minimize data movements.We present CHOPPER, a new compiler infrastructure to make Bit-serial SIMD PUD more programmable and efficient. For the better programmability, the design of CHOPPER (1) exploits bit-slicing compilers to enable automatic memory allocation and code generation, from naturally-expressive codes (i.e. similar to Parallel Haskell) into the "SIMD-Within-A-Register"-style codes; and (2) introduces a new abstraction called "Virtual Code Emitter", to make Bit-serial SIMD PUD architecture exploit Memory-Level Parallelism (i.e. Bank or Subarray) more effectively. For the better efficiency, we propose three novel optimizations for CHOPPER to better exploit the potentials of Bit-serial SIMD PUD architectures, which (1) minimize the amount of intra-subarray data movements; and (2) mitigate the overheads of spilling data outside Bit-serial SIMD PUD architectures. These optimizations can greatly improve the overall efficiency of Bit-serial SIMD PUD architectures. We also discuss (1) the limitations of the current CHOPPER; and (2) the potentials of CHOPPER for other types of Processing-In-Memory architectures.We evaluate CHOPPER by hosting it on three state-of-the-art Bit-serial SIMD PUD architectures. We compare CHOPPER-generated codes against the state-of-the-art hands-tuned codes for Bit-serial SIMD PUD architectures. We highlight that, averaged across 16 real-world workloads from 4 PUD-friendly application domains, CHOPPER achieves (A) 1.20X, 1.29X and 1.26X speedup when data can fit within DRAM subarrays; and (B) 12.61X, 9.05X and 9.81X speedup when data need to spill to the secondary storage, on Ambit [50], ELP2IM [56] and SIMDRAM [22], compared with hands-tuned codes using the state-of-the-art methodology [22] for Bit-serial SIMD PUD architectures. These performance benefits also accompany with a great reduction of Lines-of-Codes (LoC) in CHOPPER (i.e. by 4.3X less LoCs for hands-tuning a single subarray, and >103X less for hands-tuning all subarrays in a rank). We also perform breakdown and sensitivity studies of CHOPPER, to better understand its source benefits and examine its robustness under various architectural features.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Infinity Stream: Portable and Programmer-Friendly In-/Near-Memory FusionZhengrong Wang, Christopher Liu, Aman Arora, Lizy Kurian John 等ASPLOS 2023 · 被引用 20 次
- BitNN: A Bit-Serial Accelerator for K-Nearest Neighbor Search in Point CloudsMeng Han, Liang Wang, Limin Xiao, Hao Zhang 等ISCA 2024 · 被引用 14 次
- CINM (Cinnamon): A Compilation Infrastructure for Heterogeneous Compute In-Memory and Compute Near-Memory ParadigmsAsif Ali Khan, Hamid Farzaneh, Karl Friedrich Alexander Friebel, Clément Fournier 等ASPLOS 2024 · 被引用 7 次
- Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesRakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta 等HPCA 2026 · 被引用 2 次
- Count2Multiply: Reliable In-Memory High-Radix CountingJoão Paulo C. de Lima, Benjamin F. Morris III, Asif Ali Khan, Jerónimo Castrillón 等HPCA 2026
它引用的顶会 Paper10
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira 等ASPLOS 2021 · 被引用 182 次
- ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAMXin Xin, Youtao Zhang, Jun YangHPCA 2020 · 被引用 84 次
- SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory SystemsMaciej Besta, Raghavendra Kanakagiri, Grzegorz Kwasniewski, Rachata Ausavarungnirun 等MICRO 2021 · 被引用 78 次
- FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and CachingYaohua Wang, Lois Orosa, Xiangjun Peng, Yang Guo 等MICRO 2020 · 被引用 72 次
- pLUTo: Enabling Massively Parallel Computation in DRAM via Lookup TablesJoão Dinis Ferreira, Gabriel Falcão, Juan Gómez-Luna, Mohammed Alser 等MICRO 2022 · 被引用 60 次
相关 Paper
- MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data ComputingGeraldo F. Oliveira, Ataberk Olgun, Abdullah Giray Yaglikçi, F. Nisa Bostanci 等HPCA 2024 · 被引用 44 次
- UniNDP: A Unified Compilation and Simulation Tool for Near DRAM Processing ArchitecturesTongxin Xie, Zhenhua Zhu, Bing Li, Yukai He 等HPCA 2025 · 被引用 9 次
- PUMICE: Processing-using-Memory Integration with a Scalar Pipeline for Symbiotic ExecutionSocrates S. Wong, Cecilio C. Tamarit, José F. MartínezDAC 2023 · 被引用 5 次
- ATiM: Autotuning Tensor Programs for Processing-in-DRAMYongwon Shin, Dookyung Kang, Hyojin SungISCA 2025 · 被引用 2 次
- Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU CoresZhongcheng Zhang, Yan Ou, Ying Liu, Chenxi Wang 等ASPLOS 2023 · 被引用 4 次
