A Two-way SRAM Array based Accelerator for Deep Neural Network On-chip Training
Hongwu Jiang, Shanshi Huang, Xiaochen Peng, Jian-Wei Su, Yen-Chi Chou, Wei-Hsing Huang, Ta-Wei Liu, Ruhui Liu, Meng-Fan Chang, Shimeng Yu
Abstract
On-chip training of large-scale deep neural networks (DNNs) is challenging due to computational complexity and resource limitation. Compute-in-memory (CIM) architecture exploits the analog computation inside the memory array to speed up the vectormatrix multiplication (VMM) and alleviate the memory bottleneck. However, existing CIM prototype chips, in particular, SRAM-based accelerators target at implementing low-precision inference engine only. In this work, we propose a two-way SRAM array design that could perform bi-directional in-memory VMM with minimum hardware overhead. A novel solution of signed number multiplication is also proposed to handle the negative input in backpropagation. We taped-out and validated proposed two-way SRAM array design in TSMC 28nm process. Based on the silicon measurement data on CIM macro, we explore the hardware performance for the entire architecture for DNN on-chip training. The experimental data shows that proposed accelerator can achieve energy efficiency of 3.2 TOPS/W, >1000 FPS and >300 FPS for ResNet and DenseNet training on ImageNet, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 610c2be2-6eec-49c0-8b96-056aad0b90ffCited by top-tier papers1
Ask how each one uses itRelated papers
- Addition is Most You Need: Efficient Floating-Point SRAM Compute-in-Memory by Harnessing Mantissa AdditionWeidong Cao, Jian Gao, Xin Xin, Xuan ZhangDAC 2024 · 2 citations
- A Compute-in-Memory Architecture Compatible with 3D NAND Flash that Parallelly Activates Multi-LayersLiang Zhao, Chu Yan, Fan Yang, Shifan Gao et al.DAC 2021 · 15 citations
- 3D-FPIM: An Extreme Energy-Efficient DNN Acceleration System Using 3D NAND Flash-Based In-Situ PIM UnitHunjun Lee, Minseop Kim, Dongmoon Min, Joonsung Kim et al.MICRO 2022 · 23 citations
- INCA: Input-stationary Dataflow at Outside-the-box Thinking about Deep Learning AcceleratorsBokyung Kim, Shiyu Li, Hai LiHPCA 2023 · 28 citations
- Energy-efficient SNN Architecture using 3nm FinFET Multiport SRAM-based CIM with Online LearningLucas Huijbregts, Hsiao-Hsuan Liu, Paul Detterer, Said Hamdioui et al.DAC 2024 · 8 citations
