Timely: Pushing Data Movements And Interfaces In Pim Accelerators Towards Local And In Time Domain
Weitao Li, Pengfei Xu, Yang Zhao, Haitong Li, Yuan Xie, Yingyan Lin
Abstract
Resistive-random-access-memory (ReRAM) based processing-in-memory (R 2 PIM) accelerators show promise in bridging the gap between Internet of Thing devices' constrained resources and Convolutional/Deep Neural Networks' (CNNs/DNNs') prohibitive energy cost. Specifically, R 2 PIM accelerators enhance energy efficiency by eliminating the cost of weight movements and improving the computational density through ReRAM's high density. However, the energy efficiency is still limited by the dominant energy cost of input and partial sum (Psum) movements and the cost of digital-to-analog (D/A) and analog-to-digital (A/D) interfaces. In this work, we identify three energy-saving opportunities in R 2 PIM accelerators: analog data locality, time-domain interfacing, and input access reduction, and propose an innovative R 2 PIM accelerator called TIMELY, with three key contributions: (1) TIMELY adopts analog local buffers (ALBs) within ReRAM crossbars to greatly enhance the data locality, minimizing the energy overheads of both input and Psum movements; (2) TIMELY largely reduces the energy of each single D/A (and A/D) conversion and the total number of conversions by using time-domain interfaces (TDIs) and the employed ALBs, respectively; (3) we develop an only-once input read (O 2 IR) mapping method to further decrease the energy of input accesses and the number of D/A conversions. The evaluation with more than 10 CNN/DNN models and various chip configurations shows that, TIMELY outperforms the baseline R 2 PIM accelerator, PRIME, by one order of magnitude in energy efficiency while maintaining better computational density (up to 31.2×) and throughput (up to 736.6×). Furthermore, comprehensive studies are performed to evaluate the effectiveness of the proposed ALB, TDI, and O 2 IR innovations in terms of energy savings and area reduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 22da1987-b5d5-4515-9cc7-605387ffa417Cited by top-tier papers13
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li et al.NeurIPS 2020 · 99 citations
- FORMS: Fine-grained Polarized ReRAM-based In-situ Computation for Mixed-signal DNN AcceleratorGeng Yuan, Payman Behnam, Zhengang Li, Ali Shafiee et al.ISCA 2021 · 73 citations
- Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip RecomputationAmir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu KangMICRO 2022 · 47 citations
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 45 citations
- Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and AcceleratorsYonggan Fu, Yongan Zhang, Yang Zhang, David D. Cox et al.ICML 2021 · 23 citations
Related papers
- INCA: Input-stationary Dataflow at Outside-the-box Thinking about Deep Learning AcceleratorsBokyung Kim, Shiyu Li, Hai LiHPCA 2023 · 28 citations
- Optimizing ADC Utilization through Value-Aware Bypass in ReRAM-based DNN AcceleratorHanCheon Yun, Hyein Shin, Myeonggu Kang, Lee-Sup KimDAC 2021 · 5 citations
- Towards State-Aware Computation in ReRAM Neural NetworksYintao He, Ying Wang, Xiandong Zhao, Huawei Li et al.DAC 2020 · 8 citations
- InfoX: an energy-efficient ReRAM accelerator design with information-lossless low-bit ADCsYintao He, Songyun Qu, Ying Wang, Bing Li et al.DAC 2022 · 10 citations
- F3D: Accelerating 3D Convolutional Neural Networks in Frequency Space Using ReRAMBosheng Liu, Zhuoshen Jiang, Jigang Wu, Xiaoming Chen et al.DAC 2021 · 4 citations
