MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural Networks
Seah Kim, Hasan Genc, Vadim Vadimovich Nikiforov, Krste Asanovic, Borivoje Nikolic, Yakun Sophia Shao
摘要
Driven by the wide adoption of deep neural networks (DNNs) across different application domains, multitenancy execution, where multiple DNNs are deployed simultaneously on the same hardware, has been proposed to satisfy the latency requirements of different applications while improving the overall system utilization. However, multi-tenancy execution could lead to undesired system-level resource contention, causing quality-of-service (QoS) degradation for latency-critical applications.
To address this challenge, we propose MOCA 1 , an adaptive multi-tenancy system for DNN accelerators. Unlike existing solutions that focus on compute resource partition, MOCA dynamically manages shared memory resources of co-located applications to meet their QoS targets. Specifically, MOCA leverages the regularities in both DNN operators and accelerators to dynamically modulate memory access rates based on their latency targets and user-defined priorities so that co-located applications get the resources they demand without significantly starving their co-runners. We demonstrate that MOCA improves the satisfaction rate of the service level agreement (SLA) up to 3.9× (1.8× average), system throughput by 2.3× (1.7× average), and fairness by 1.3× (1.2× average), compared to prior work.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- DREAM: A Dynamic Scheduler for Dynamic Real-time Multi-model ML WorkloadsSeah Kim, Hyoukjun Kwon, Jinook Song, Jihyuck Jo 等ASPLOS 2023 · 被引用 20 次
- Shared Memory-contention-aware Concurrent DNN Execution for Diversely Heterogeneous System-on-ChipsIsmet Dagli, Mehmet E. BelviranliPPoPP 2024 · 被引用 18 次
- RoSÉ: A Hardware-Software Co-Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC EvaluationDima Nikiforov, Shengjun Chris Dong, Chengyi Lux Zhang, Seah Kim 等ISCA 2023 · 被引用 17 次
- sNPU: Trusted Execution Environments on Integrated NPUsErhu Feng, Dahu Feng, Dong Du, Yubin Xia 等ISCA 2024 · 被引用 13 次
- SCAR: Scheduling Multi-Model AI Workloads on Heterogeneous Multi-Chiplet Module AcceleratorsMohanad Odema, Luke Chen, Hyoukjun Kwon, Mohammad Abdullah Al FaruqueMICRO 2024 · 被引用 11 次
它引用的顶会 Paper9
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson 等ISCA 2020 · 被引用 517 次
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali 等DAC 2021 · 被引用 325 次
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 被引用 150 次
- Heterogeneous Dataflow Accelerators for Multi-DNN WorkloadsHyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna 等HPCA 2021 · 被引用 143 次
- Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural NetworksSoroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer 等MICRO 2020 · 被引用 120 次
相关 Paper
- AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant WorkloadsSeah Kim, Jerry Zhao, Krste Asanovic, Borivoje Nikolic 等MICRO 2023 · 被引用 10 次
- CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUsTianhao Cai, Liang Wang, Limin Xiao, Meng Han 等DAC 2025
- VELTAIR: towards high-performance multi-tenant deep learning services via adaptive compilation and schedulingZihan Liu, Jingwen Leng, Zhihui Zhang, Quan Chen 等ASPLOS 2022 · 被引用 52 次
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 被引用 110 次
- A Deep Reinforcement Learning based Online Scheduling Policy for Deep Neural Network Multi-Tenant Multi-Accelerator SystemsFrancesco Giulio Blanco, Enrico Russo, Maurizio Palesi, Davide Patti 等DAC 2024 · 被引用 9 次
