CoExe: An Efficient Co-execution Architecture for Real-Time Neural Network Services
Chubo Liu, Kenli Li, Mingcong Song, Jiechen Zhao, Keqin Li, Tao Li, Zihao Zeng
Abstract
End-to-end latency is sensitive for user-interactive neural network (NN) services on clouds. For periods of high request load, co-locating multiple NN requests has the potential to reduce end-to-end latency. However, current batch-based accelerators lack request-level parallelism support, leaving the queuing time non-optimized. Meanwhile, naively partitioning resources for simultaneous requests suffers from longer execution time as well as lower resource efficiency because different applications utilize separate resources without sharing. To effectively reduce the end-to-end latency for real-time NN requests, we propose CoExe architecture, equipped with a pipeline implementation of a sparsity-driven real-time co-execution model. By leveraging the non-trivial amount of sparse operations during concurrent NNs execution, the end-to-end latency is decreased by up to 12.3× and 2.4× over Eyeriss-like and SCNN at peak workload mode. Besides, we propose row cross (RC) dataflow to reduce data movement cost, and avoid memory duplication.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0d554e08-c217-4770-8da6-71c439e87c61Related papers
- ISOSceles: Accelerating Sparse CNNs through Inter-Layer PipeliningYifan Yang, Joel S. Emer, Daniel SánchezHPCA 2023 · 27 citations
- SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural NetworksSurya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin et al.ISCA 2020 · 122 citations
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis et al.MobiCom 2020 · 312 citations
- Real-time neural network inference on extremely weak devices: agile offloading with explainable AIKai Huang, Wei GaoMobiCom 2022 · 57 citations
- DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in VideosMathias Parger, Chengcheng Tang, Christopher D. Twigg, Cem Keskin et al.CVPR 2022 · 31 citations
