Addition is Most You Need: Efficient Floating-Point SRAM Compute-in-Memory by Harnessing Mantissa Addition
Weidong Cao, Jian Gao, Xin Xin, Xuan Zhang
Abstract
The compute-in-memory (CIM) paradigm holds great promise to efficiently accelerate machine learning workloads. Among memory devices, static random-access memory (SRAM) stands out as a practical choice for its exceptional reliability in the digital domain and excellent scalability. Recently, there has been a growing interest in accelerating floating-point (FP) deep neural networks (DNNs) with SRAM CIM due to their critical importance in DNN training and high-accurate inference. This paper proposes an energy-efficient SRAM CIM macro for FP DNNs. To achieve the design, we identify a lightweight approach that decomposes conventional FP mantissa multiplication into two parts: mantissa sub-addition (sub-ADD) and mantissa sub-multiplication (sub-MUL). Our study shows that while mantissa sub-MUL is compute-intensive, it only contributes to the minority of FP products, whereas mantissa sub-ADD, although compute-light, accounts for the majority of FP products. Recognizing "Addition is Most You Need", we develop a novel hybrid-domain SRAM CIM macro to accurately handle mantissa sub-ADD in the digital domain while improving the energy efficiency of mantissa sub-MUL using analog computing. Experiments with the MLPerf benchmark show its remarkable improvement in energy efficiency on average by 3× 3.6× (2.5× 3.1×) in inference (training) compared to a fully digital baseline without any accuracy loss, showcasing its great potential for FP DNN acceleration.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Improving the Efficiency of In-Memory-Computing Macro with a Hybrid Analog-Digital Computing Mode for Lossless Neural Network InferenceQilin Zheng, Ziru Li, Jonathan Ku, Yitu Wang et al.DAC 2024 · 2 citations
- A Two-way SRAM Array based Accelerator for Deep Neural Network On-chip TrainingHongwu Jiang, Shanshi Huang, Xiaochen Peng, Jian-Wei Su et al.DAC 2020 · 39 citations
- Efficient Memory Integration: MRAM-SRAM Hybrid Accelerator for Sparse On-Device LearningFan Zhang, Amitesh Sridharan, Wilman Tsai, Yiran Chen et al.DAC 2024 · 7 citations
- Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN InferenceHongrui Guo, Mo Zou, Yifan Hao, Zidong Du et al.MICRO 2024 · 3 citations
- Energy Efficient Dual Designs of FeFET-Based Analog In-Memory Computing with Inherent Shift-Add CapabilityZeyu Yang, Qingrong Huang, Yu Qian, Kai Ni et al.DAC 2024 · 9 citations
