Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks
Soroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer, Brahmendra Reddy Yatham, Navateja Alla, Hardik Sharma, Mohammad Alian, Eiman Ebrahimi, Nam Sung Kim, Cliff Young, Hadi Esmaeilzadeh
Abstract
Deep Neural Networks (DNNs) have reinvigorated real-world applications that rely on learning patterns of data and are permeating into different industries and markets. Cloud infrastructure and accelerators that offer INFerence-as-a-Service (INFaaS) have become the enabler of this rather quick and invasive shift in the industry. To that end, mostly accelerator-based INFaaS (Google's TPU [1], NVIDIA T4 [2], Microsoft Brainwave [3], etc.) has become the backbone of many real-life applications. However, as the demand for such services grows, merely scaling-out the number of accelerators is not economically cost-effective. Although multi-tenancy has propelled datacenter scalability, it has not been a primary factor in designing DNN accelerators due to the arms race for higher speed and efficiency. This paper sets out to explore this timely requirement of multi-tenancy through a new dimension: dynamic architecture fission. To that end, we define Planaria1that can dynamically fission (break) into multiple smaller yet full-fledged DNN engines at runtime. This microarchitectural capability enables spatially co-locating multiple DNN inference services on the same hardware, offering simultaneous multi-tenant DNN acceleration. To realize this dynamic reconfigurability, we first devise breakable omni-directional systolic arrays for DNN acceleration that allows omni-directional flow of data. Second, it uses this capability and a unique organization of on-chip memory, interconnection, and compute resources to enable fission in systolic array based DNN accelerators. Architecture fission and its associated flexibility enables an extra degree of freedom for task scheduling, that even allows breaking the accelerator with regard to the server load, DNN topology, and task priority. As such, it can simultaneously co-locate DNNs to enhance utilization, throughput, QoS, and fairness. We compare the proposed design to PREMA [4], a recent effort that offers multi-tenancy by time-multiplexing the DNN accelerator across multiple tasks. We use the same frequency, the same amount of compute and memory resources for both accelerators. The results show significant benefits with (soft, medium, hard) QoS requirements, in throughput (7.4×, 7.2×, 12.2×), SLA satisfaction rate (45%, 15%, 16%), and fairness (2.1×, 2.3×, 1.9×).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers31
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- Transparent GPU Sharing in Container Clouds for Deep Learning WorkloadsBingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu et al.NSDI 2023 · 112 citations
- MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator CoresSheng-Chun Kao, Tushar KrishnaHPCA 2022 · 58 citations
- VELTAIR: towards high-performance multi-tenant deep learning services via adaptive compilation and schedulingZihan Liu, Jingwen Leng, Zhihui Zhang, Quan Chen et al.ASPLOS 2022 · 52 citations
Builds on5
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 150 citations
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 110 citations
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network CompilationByung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, Hadi EsmaeilzadehICLR 2020 · 90 citations
- Bit-Parallel Vector Composability for Neural AccelerationSoroush Ghodrati, Hardik Sharma, Cliff Young, Nam Sung Kim et al.DAC 2020 · 19 citations
Related papers
- MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural NetworksSeah Kim, Hasan Genc, Vadim Vadimovich Nikiforov, Krste Asanovic et al.HPCA 2023 · 36 citations
- Adyna: Accelerating Dynamic Neural Networks with Adaptive SchedulingZhiyao Li, Bohan Yang, Jiaxiang Li, Taijie Chen et al.HPCA 2025 · 2 citations
- AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant WorkloadsSeah Kim, Jerry Zhao, Krste Asanovic, Borivoje Nikolic et al.MICRO 2023 · 10 citations
- Partitioned Scheduling and Parallelism Assignment for Real-Time DNN Inference Tasks on Multi-TPUBinqi Sun, Tomasz Kloda, Chu-Ge Wu, Marco CaccamoDAC 2024 · 8 citations
- Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice SparsityDongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong et al.HPCA 2025 · 5 citations
