Count2Multiply: Reliable In-Memory High-Radix Counting
João Paulo C. de Lima, Benjamin F. Morris III, Asif Ali Khan, Jerónimo Castrillón, Alex K. Jones
Abstract
Computing-in-memory (CIM) has been demonstrated across various memory technologies, ranging from memristive crossbars performing analog dot-product computations to large-scale digital bitwise operations in commodity DRAM and other proposed nonvolative memory technologies. However, current CIM solutions face latency and reliability challenges. CIM fidelity lags considerably behind access fidelity. Furthermore, bulk-bitwise CIM, although highly parallelized, requires long latency for operations like multiplication and addition, due to their bit-serial computation.
This paper presents Count2Multiply, a technology-agnostic digital CIM approach to perform multiplication, addition and other operations using high-radix, massively parallel counting enabled by CIM bulk-bitwise logic operations. Designed to meet fault tolerance requirements, Count2Multiply integrates traditional row-wise error correction codes, such as Hamming and BCH, to address the high error rates in existing CIM designs. We demonstrate Count2Multiply with a detailed application to CIM in conventional DRAM due to its ubiquity and high endurance. However, we note that the Count2Multiply architecture is compatible with other functionally complete CIM proposals. Compared to the state-of-the-art in-DRAM CIM method, Count2Multiply achieves up to 10× speedup, 8× higher GOPS/Watt, and 9.5× higher GOPS/area, while outperforming GPU for vector-matrix multiplications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57415248-10ca-4ed9-83d2-6384fdcc2794Cited by top-tier papers1
Ask how each one uses itBuilds on12
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
- SIMDRAM: a framework for bit-serial SIMD processing using DRAMNastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, João Dinis Ferreira et al.ASPLOS 2021 · 182 citations
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li et al.NeurIPS 2020 · 99 citations
- ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAMXin Xin, Youtao Zhang, Jun YangHPCA 2020 · 84 citations
- MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data ComputingGeraldo F. Oliveira, Ataberk Olgun, Abdullah Giray Yaglikçi, F. Nisa Bostanci et al.HPCA 2024 · 44 citations
Related papers
- ER-DCIM: Error-Resilient Digital CIM Architecture with Run-Time MAC-Cell Error CorrectionZhen He, Yiqi Wang, Zihan Wu, Shaojun Wei et al.HPCA 2025 · 3 citations
- HR-DCIM: High-Reliability Floating-Point Digital CIM Architecture With Unified Low-Cost Iterative Error CorrectionZhen He, Yiqi Wang, Zhiheng Yue, Zihan Wu et al.HPCA 2026
- Improving compute in-memory ECC reliability with successive correctionBrian Crafton, Zishen Wan, Samuel Spetalnick, Jong-Hyeok Yoon et al.DAC 2022 · 17 citations
- CorcPUM: Efficient Processing Using Cross-Point Memory via Cooperative Row-Column Access Pipelining and Adaptive Timing Optimization in SubarraysChengning Wang, Dan Feng, Wei Tong, Jingning LiuDAC 2023 · 3 citations
- A Two-way SRAM Array based Accelerator for Deep Neural Network On-chip TrainingHongwu Jiang, Shanshi Huang, Xiaochen Peng, Jian-Wei Su et al.DAC 2020 · 39 citations
