DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow Accelerators
Xiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong, Guilherme Paim, Marian Verhelst
摘要
Deep Neural Networks (DNNs) have achieved remarkable success across various intelligent tasks but encounter performance and energy challenges in inference execution due to data movement bottlenecks. We introduce DataMaestro, a versatile and efficient data streaming unit that brings the decoupled access/execute architecture to DNN dataflow accelerators to address this issue. DataMaestro supports flexible and programmable access patterns to accommodate diverse workload types and dataflows, incorporates fine-grained prefetch and addressing mode switching to mitigate bank conflicts, and enables customizable on-the-fly data manipulation to reduce memory footprints and access counts. We integrate five DataMaestros with a Tensor Core-like GeMM accelerator and a Quantization accelerator into a RISC-V host system for evaluation. The FPGA prototype and VLSI synthesis results demonstrate that DataMaestro helps the GeMM core achieve nearly 100% utilization, which is 1.05 better than state-of-the-art solutions, while minimizing area and energy consumption to merely 6.43% and 15.06% of the total system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali 等DAC 2021 · 被引用 325 次
- HybridDNN: A Framework for High-Performance Hybrid DNN Accelerator Design and ImplementationHanchen Ye, Xiaofan Zhang, Zhize Huang, Gengsheng Chen 等DAC 2020 · 被引用 72 次
- BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning AccelerationMan Shi, Vikram Jain, Antony Joseph, Maurice Meijer 等HPCA 2024 · 被引用 46 次
- FEATHER: A Reconfigurable Accelerator with Data Reordering Support for Low-Cost On-Chip Dataflow SwitchingJianming Tong, Anirudh Itagi, Prasanth Chatarasi, Tushar KrishnaISCA 2024 · 被引用 36 次
相关 Paper
- Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesEnrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó 等HPCA 2023 · 被引用 31 次
- Duplo: Lifting Redundant Memory Accesses of Deep Neural Networks for GPU Tensor CoresHyeonjin Kim, Sungwoo Ahn, Yunho Oh, Bogil Kim 等MICRO 2020 · 被引用 27 次
- Effectively Scheduling Computational Graphs of Deep Neural Networks toward Their Domain-Specific AcceleratorsJie Zhao, Siyuan Feng, Xiaoqiang Dan, Fei Liu 等OSDI 2023 · 被引用 9 次
- GPNPU: Enabling Efficient Hardware-Based Direct Convolution with Multi-Precision Support in GPU Tensor CoresZhuoran Song, Jianfei Wang, Tianjian Li, Li Jiang 等DAC 2020 · 被引用 12 次
- BiQGEMM: matrix multiplication with lookup table for binary-coding-based quantized DNNsYongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim 等SC 2020 · 被引用 31 次
