MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural Networks
Seah Kim, Hasan Genc, Vadim Vadimovich Nikiforov, Krste Asanovic, Borivoje Nikolic, Yakun Sophia Shao
Abstract
Driven by the wide adoption of deep neural networks (DNNs) across different application domains, multitenancy execution, where multiple DNNs are deployed simultaneously on the same hardware, has been proposed to satisfy the latency requirements of different applications while improving the overall system utilization. However, multi-tenancy execution could lead to undesired system-level resource contention, causing quality-of-service (QoS) degradation for latency-critical applications.
To address this challenge, we propose MOCA 1 , an adaptive multi-tenancy system for DNN accelerators. Unlike existing solutions that focus on compute resource partition, MOCA dynamically manages shared memory resources of co-located applications to meet their QoS targets. Specifically, MOCA leverages the regularities in both DNN operators and accelerators to dynamically modulate memory access rates based on their latency targets and user-defined priorities so that co-located applications get the resources they demand without significantly starving their co-runners. We demonstrate that MOCA improves the satisfaction rate of the service level agreement (SLA) up to 3.9× (1.8× average), system throughput by 2.3× (1.7× average), and fairness by 1.3× (1.2× average), compared to prior work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d86fcd35-83ba-4fae-bf18-eba6dd2b5d9eCited by top-tier papers10
- DREAM: A Dynamic Scheduler for Dynamic Real-time Multi-model ML WorkloadsSeah Kim, Hyoukjun Kwon, Jinook Song, Jihyuck Jo et al.ASPLOS 2023 · 20 citations
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsIsmet Dagli, Mehmet E. BelviranliPPoPP 2024 · 18 citations
- RoSÉ: A Hardware-Software Co-Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC EvaluationDima Nikiforov, Shengjun Chris Dong, Chengyi Lux Zhang, Seah Kim et al.ISCA 2023 · 17 citations
- sNPU: Trusted Execution Environments on Integrated NPUsErhu Feng, Dahu Feng, Dong Du, Yubin Xia et al.ISCA 2024 · 13 citations
- SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module AcceleratorsMohanad Odema, Luke Chen, Hyoukjun Kwon, Mohammad Abdullah Al FaruqueMICRO 2024 · 11 citations
Builds on9
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali et al.DAC 2021 · 325 citations
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 150 citations
- Heterogeneous Dataflow Accelerators for Multi-DNN WorkloadsHyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna et al.HPCA 2021 · 143 citations
- Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural NetworksSoroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer et al.MICRO 2020 · 120 citations
Related papers
- AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant WorkloadsSeah Kim, Jerry Zhao, Krste Asanovic, Borivoje Nikolic et al.MICRO 2023 · 10 citations
- CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsTianhao Cai, Liang Wang, Limin Xiao, Meng Han et al.DAC 2025
- VELTAIR: towards high-performance multi-tenant deep learning services via adaptive compilation and schedulingZihan Liu, Jingwen Leng, Zhihui Zhang, Quan Chen et al.ASPLOS 2022 · 52 citations
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 110 citations
- A Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator SystemsFrancesco Giulio Blanco, Enrico Russo, Maurizio Palesi, Davide Patti et al.DAC 2024 · 9 citations
