Timely: Pushing Data Movements And Interfaces In Pim Accelerators Towards Local And In Time Domain
Weitao Li, Pengfei Xu, Yang Zhao, Haitong Li, Yuan Xie, Yingyan Lin
摘要
Resistive-random-access-memory (ReRAM) based processing-in-memory (R 2 PIM) accelerators show promise in bridging the gap between Internet of Thing devices' constrained resources and Convolutional/Deep Neural Networks' (CNNs/DNNs') prohibitive energy cost. Specifically, R 2 PIM accelerators enhance energy efficiency by eliminating the cost of weight movements and improving the computational density through ReRAM's high density. However, the energy efficiency is still limited by the dominant energy cost of input and partial sum (Psum) movements and the cost of digital-to-analog (D/A) and analog-to-digital (A/D) interfaces. In this work, we identify three energy-saving opportunities in R 2 PIM accelerators: analog data locality, time-domain interfacing, and input access reduction, and propose an innovative R 2 PIM accelerator called TIMELY, with three key contributions: (1) TIMELY adopts analog local buffers (ALBs) within ReRAM crossbars to greatly enhance the data locality, minimizing the energy overheads of both input and Psum movements; (2) TIMELY largely reduces the energy of each single D/A (and A/D) conversion and the total number of conversions by using time-domain interfaces (TDIs) and the employed ALBs, respectively; (3) we develop an only-once input read (O 2 IR) mapping method to further decrease the energy of input accesses and the number of D/A conversions. The evaluation with more than 10 CNN/DNN models and various chip configurations shows that, TIMELY outperforms the baseline R 2 PIM accelerator, PRIME, by one order of magnitude in energy efficiency while maintaining better computational density (up to 31.2×) and throughput (up to 736.6×). Furthermore, comprehensive studies are performed to evaluate the effectiveness of the proposed ALB, TDI, and O 2 IR innovations in terms of energy savings and area reduction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- ShiftAddNet: A Hardware-Inspired Deep NetworkHaoran You, Xiaohan Chen, Yongan Zhang, Chaojian Li 等NeurIPS 2020 · 被引用 99 次
- FORMS: Fine-grained Polarized ReRAM-based In-situ Computation for Mixed-signal DNN AcceleratorGeng Yuan, Payman Behnam, Zhengang Li, Ali Shafiee 等ISCA 2021 · 被引用 73 次
- Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip RecomputationAmir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu KangMICRO 2022 · 被引用 47 次
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 被引用 45 次
- Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and AcceleratorsYonggan Fu, Yongan Zhang, Yang Zhang, David D. Cox 等ICML 2021 · 被引用 23 次
相关 Paper
- INCA: Input-stationary Dataflow at Outside-the-box Thinking about Deep Learning AcceleratorsBokyung Kim, Shiyu Li, Hai LiHPCA 2023 · 被引用 28 次
- Optimizing ADC Utilization through Value-Aware Bypass in ReRAM-based DNN AcceleratorHanCheon Yun, Hyein Shin, Myeonggu Kang, Lee-Sup KimDAC 2021 · 被引用 5 次
- Towards State-Aware Computation in ReRAM Neural NetworksYintao He, Ying Wang, Xiandong Zhao, Huawei Li 等DAC 2020 · 被引用 8 次
- InfoX: an energy-efficient ReRAM accelerator design with information-lossless low-bit ADCsYintao He, Songyun Qu, Ying Wang, Bing Li 等DAC 2022 · 被引用 10 次
- F3D: Accelerating 3D Convolutional Neural Networks in Frequency Space Using ReRAMBosheng Liu, Zhuoshen Jiang, Jigang Wu, Xiaoming Chen 等DAC 2021 · 被引用 4 次
