MAPA: multi-accelerator pattern allocation policy for multi-tenant GPU servers
Kiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano, Shuaiwen Leon Song, Daniel Wong
摘要
Multi-accelerator servers are increasingly being deployed in shared multi-tenant environments (such as in cloud data centers) in order to meet the demands of large-scale compute-intensive workloads. In addition, these accelerators are increasingly being inter-connected in complex topologies and workloads are exhibiting a wider variety of inter-accelerator communication patterns. However, existing allocation policies are ill-suited for these emerging use-cases. Specifically, this work identifies that multi-accelerator workloads are commonly fragmented leading to reduced bandwidth and increased latency for inter-accelerator communication.
We propose Multi-Accelerator Pattern Allocation (MAPA), a graph pattern mining approach towards providing generalized allocation support for allocating multi-accelerator workloads on multi-accelerator servers. We demonstrate that MAPA is able to improve the execution time of multi-accelerator workloads and that MAPA is able to provide generalized benefits across various accelerator topologies. Finally, we demonstrate a speedup of 12.4% for 75th percentile of jobs with the worst case execution time reduced by up to 35% against baseline policy using MAPA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- Hi-Speed DNN Training with Espresso: Unleashing the Full Potential of Gradient Compression with Near-Optimal Usage StrategiesZhuang Wang, Haibin Lin, Yibo Zhu, T. S. Eugene NgEuroSys 2023 · 被引用 26 次
- Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric ApproachHao Wu, Yaochen Liu, Minchen Yu, Qizhen Weng 等EuroSys 2026
它引用的顶会 Paper5
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Peregrine: a pattern-aware graph mining systemKasra Jamshidi, Rakesh Mahadasa, Keval VoraEuroSys 2020 · 被引用 107 次
- Smart-PGSim: using neural network to accelerate AC-OPF power grid simulationWenqian Dong, Zhen Xie, Gokcen Kestor, Dong LiSC 2020 · 被引用 44 次
- BlockMaestro: Enabling Programmer-Transparent Task-based Execution in GPU SystemsAmirAli Abdolrashidi, Hodjat Asghari Esfeden, Ali Jahanshahi, Kaustubh Singh 等ISCA 2021 · 被引用 15 次
- Independent Forward Progress of Work-groupsAlexandru Dutu, Matthew D. Sinclair, Bradford M. Beckmann, David A. Wood 等ISCA 2020 · 被引用 5 次
相关 Paper
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator SystemsGuan Shen, Jieru Zhao, Zeke Wang, Zhe Lin 等DAC 2023 · 被引用 5 次
- FlexMiner: A Pattern-Aware Accelerator for Graph Pattern MiningXuhao Chen, Tianhao Huang, Shuotao Xu, Thomas Bourgeat 等ISCA 2021 · 被引用 41 次
- AvA: Accelerated Virtualization of AcceleratorsHangchen Yu, Arthur Michener Peters, Amogh Akshintala, Christopher J. RossbachASPLOS 2020 · 被引用 33 次
- MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator CoresSheng-Chun Kao, Tushar KrishnaHPCA 2022 · 被引用 58 次
- LCCG: a locality-centric hardware accelerator for high throughput of concurrent graph processingJin Zhao, Yu Zhang, Xiaofei Liao, Ligang He 等SC 2021 · 被引用 8 次
