SC2021Top-tier venue
MAPA: multi-accelerator pattern allocation policy for multi-tenant GPU servers
Kiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano, Shuaiwen Leon Song, Daniel Wong
Abstract
Multi-accelerator servers are increasingly being deployed in shared multi-tenant environments (such as in cloud data centers) in order to meet the demands of large-scale compute-intensive workloads. In addition, these accelerators are increasingly being inter-connected in complex topologies and workloads are exhibiting a wider variety of inter-accelerator communication patterns. However, existing allocation policies are ill-suited for these emerging use-cases. Specifically, this work identifies that multi-accelerator workloads are commonly fragmented leading to reduced bandwidth and increased latency for inter-accelerator communication.
We propose Multi-Accelerator Pattern Allocation (MAPA), a graph pattern mining approach towards providing generalized allocation support for allocating multi-accelerator workloads on multi-accelerator servers. We demonstrate that MAPA is able to improve the execution time of multi-accelerator workloads and that MAPA is able to provide generalized benefits across various accelerator topologies. Finally, we demonstrate a speedup of 12.4% for 75th percentile of jobs with the worst case execution time reduced by up to 35% against baseline policy using MAPA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fab6930-4053-489c-bfd0-efc14f5b8387Cited by top-tier papers3
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- Hi-Speed DNN Training with Espresso: Unleashing the Full Potential of Gradient Compression with Near-Optimal Usage StrategiesZhuang Wang, Haibin Lin, Yibo Zhu, T. S. Eugene NgEuroSys 2023 · 26 citations
- Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric ApproachHao Wu, Yaochen Liu, Minchen Yu, Qizhen Weng et al.EuroSys 2026
Builds on5
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Peregrine: a pattern-aware graph mining systemKasra Jamshidi, Rakesh Mahadasa, Keval VoraEuroSys 2020 · 107 citations
- Smart-PGSim: using neural network to accelerate AC-OPF power grid simulationWenqian Dong, Zhen Xie, Gokcen Kestor, Dong LiSC 2020 · 44 citations
- BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU SystemsAmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh et al.ISCA 2021 · 15 citations
- Independent Forward Progress of Work-groupsAlexandru Dutu, Matthew D. Sinclair, Bradford M. Beckmann, David A. Wood et al.ISCA 2020 · 5 citations
Related papers
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin et al.DAC 2023 · 5 citations
- FlexMiner: A Pattern-Aware Accelerator for Graph Pattern MiningXuhao Chen, Tianhao Huang, Shuotao Xu, Thomas Bourgeat et al.ISCA 2021 · 41 citations
- AvA: Accelerated Virtualization of AcceleratorsHangchen Yu, Arthur Michener Peters, Amogh Akshintala, Christopher J. RossbachASPLOS 2020 · 33 citations
- MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator CoresSheng-Chun Kao, Tushar KrishnaHPCA 2022 · 58 citations
- LCCG: a locality-centric hardware accelerator for high throughput of concurrent graph processingJin Zhao, Yu Zhang, Xiaofei Liao, Ligang He et al.SC 2021 · 8 citations
