DeepBurning-SEG: Generating DNN Accelerators of Segment-Grained Pipeline Architecture
Xuyi Cai, Ying Wang, Xiaohan Ma, Yinhe Han, Lei Zhang
摘要
The growing complexity and diversity of deep neural network (DNN) applications have inspired intensive research on specialized DNN accelerators and also the design automation frameworks. Previous specialized NN acceleratos roughly fall into two categories of implementation, either the no-pipelined architecture that relies on a generic processing unit (PU) to sequentially execute the DNN layers in a layer-wise way, or the fully-pipelined architecture that dedicates interconnected customized PUs to the corresponding DNN layers in the model. Thus, such designs often suffer from either the resource under-utilization issue faced by no-pipelined accelerators or the resource scalability problem brought by the over-deep pipeline designs. In this work, we propose a novel class of design solution for DNN acceleration, segment-grained pipeline architecture (SPA). In the SPA accelerator, the targeted workload of DNN models will be divided into many segments and each segment will be sequentially executed on the shared interconnected PUs in a pipeline manner, so that they will benefit from both the efficiency of pipelined execution and also the flexibility of sharing PUs across different model layers. Particularly, we found that the efficiency of the implemented SPA accelerator significantly depends on the segmentation strategies of the models and the hardware resources assignment policy for PUs. Therefore, we introduce an automated design framework, AutoSeg, that includes a parameterized SPA accelerator template and a co-design engine that will generate the efficient model segmentation solution and hardware pipeline design parameters for the acceleration workload. Experimental results show that the SPA solutions generated by the AutoSeg framework achieve speedup when compared to ASIC-based general DNN processors, and the FPGA designs implemented by AutoSeg also achieve as high as DSP efficiency and throughput improvement.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingHuizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue 等MICRO 2024 · 被引用 31 次
- Reconfigurable Stream Network ArchitectureChengyue Wang, Xiaofan Zhang, Jason Cong, James C. HoeISCA 2025 · 被引用 8 次
- LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning AcceleratorGuoyu Li, Shengyu Ye, Chunyun Chen, Yang Wang 等HPCA 2025 · 被引用 7 次
- GoPIM: GCN-Oriented Pipeline Optimization for PIM AcceleratorsSiling Yang, Shuibing He, Wenjiong Wang, Yanlong Yin 等HPCA 2025 · 被引用 3 次
- Empowering Vector Architectures for ML: The CAMP Architecture for Matrix MultiplicationMohammadreza Esmali Nojehdeh, Hossein Mokhtarnia, Julian Pavon, Narcís Rodas 等MICRO 2025 · 被引用 1 次
相关 Paper
- QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioningJongho Park, Hyukjun Kwon, Seowoo Kim, Junyoung Lee 等DAC 2022 · 被引用 6 次
- Auto-NBA: Efficient and Effective Search Over the Joint Space of Networks, Bitwidths, and AcceleratorsYonggan Fu, Yongan Zhang, Yang Zhang, David D. Cox 等ICML 2021 · 被引用 23 次
- Deep Learning Acceleration with Neuron-to-Memory TransformationMohsen Imani, Mohammad Samragh Razlighi, Yeseong Kim, Saransh Gupta 等HPCA 2020 · 被引用 31 次
- HybridDNN: A Framework for High-Performance Hybrid DNN Accelerator Design and ImplementationHanchen Ye, Xiaofan Zhang, Zhize Huang, Gengsheng Chen 等DAC 2020 · 被引用 72 次
- Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple TasksLei Yang, Zheyu Yan, Meng Li, Hyoukjun Kwon 等DAC 2020 · 被引用 115 次
