AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant Workloads
Seah Kim, Jerry Zhao, Krste Asanovic, Borivoje Nikolic, Yakun Sophia Shao
Abstract
With the widespread adoption of deep neural networks (DNNs) across applications, there is a growing demand for DNN deployment solutions that can seamlessly support multi-tenant execution. This involves simultaneously running multiple DNN workloads on heterogeneous architectures with domain-specific accelerators. However, existing accelerator interfaces directly bind the accelerator's physical resources to user threads, without an efficient mechanism to adaptively re-partition available resources. This leads to high programming complexities and performance overheads due to sub-optimal resource allocation, making scalable many-accelerator deployment impractical.
To address this challenge, we propose AuRORA, a novel accelerator integration methodology that enables scalable accelerator deployment for multi-tenant workloads. In particular, AuRORA supports virtualized accelerator orchestration via co-designing the hardware-software stack of accelerators to allow adaptively binding current workloads onto available accelerators. We demonstrate that AuRORA achieves 2.02× higher overall SLA satisfaction, 1.33× overall system throughput, and 1.34× overall fairness compared to existing accelerator integration solutions with less than 2.7% area overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 95bbd7ec-3c05-4b61-ac89-2ba02ec2aae0Cited by top-tier papers6
- sNPU: Trusted Execution Environments on Integrated NPUsErhu Feng, Dahu Feng, Dong Du, Yubin Xia et al.ISCA 2024 · 13 citations
- SuperNoVA: Algorithm-Hardware Co-Design for Resource-Aware SLAMSeah Kim, Roger Hsiao, Borivoje Nikolic, James Demmel et al.ASPLOS 2025 · 7 citations
- Practical Federated Recommendation Model Learning Using ORAM with Controlled PrivacyJinyu Liu, Wenjie Xiong, G. Edward Suh, Kiwan MaengASPLOS 2025 · 2 citations
- Topology-Aware Virtualization over Inter-Core Connected Neural Processing UnitsDahu Feng, Erhu Feng, Dong Du, Pinjie Xu et al.ISCA 2025 · 2 citations
- AccelFlow: Orchestrating an On-Package Ensemble of Fine-Grained Accelerators for MicroservicesJovan Stojkovic, Abraham Farrell, Zhangxiaowen Gong, Christopher J. Hughes et al.HPCA 2026 · 2 citations
Builds on10
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationHasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali et al.DAC 2021 · 325 citations
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 150 citations
- Heterogeneous Dataflow Accelerators for Multi-DNN WorkloadsHyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna et al.HPCA 2021 · 143 citations
- Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural NetworksSoroush Ghodrati, Byung Hoon Ahn, Joon Kyung Kim, Sean Kinzer et al.MICRO 2020 · 120 citations
Related papers
- MoCA: Memory-Centric, Adaptive Execution for Multi-Tenant Deep Neural NetworksSeah Kim, Hasan Genc, Vadim Vadimovich Nikiforov, Krste Asanovic et al.HPCA 2023 · 36 citations
- Atomic Dataflow based Graph-Level Workload Orchestration for Scalable DNN AcceleratorsShixuan Zheng, Xianjue Zhang, Leibo Liu, Shaojun Wei et al.HPCA 2022 · 39 citations
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin et al.DAC 2023 · 5 citations
- A Multi-Neural Network Acceleration ArchitectureEunjin Baek, Dongup Kwon, Jangwoo KimISCA 2020 · 110 citations
- Adyna: Accelerating Dynamic Neural Networks with Adaptive SchedulingZhiyao Li, Bohan Yang, Jiaxiang Li, Taijie Chen et al.HPCA 2025 · 2 citations
