USENIX ATC2024顶会
Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma, Aleksandar Kuzmanovic
摘要
Modern machine learning (ML) workloads heavily depend on distributing tasks across clusters of server CPUs and specialized accelerators, such as GPUs and TPUs, to achieve optimal performance. Nonetheless, prior research has highlighted the inefficient utilization of computing resources in distributed ML, leading to suboptimal performance. This inefficiency primarily stems from CPU bottlenecks and suboptimal accelerator scheduling. Although numerous proposals have been put forward to address these issues individually, none have effectively tackled both inefficiencies simultaneously. In this paper, we introduce Conspirator, an innovative control plane design aimed at alleviating both bottlenecks by harnessing the enhanced computing capabilities of SmartNICs. Following the evolving role of SmartNICs, which have transitioned from their initial function of standard networking task offloading to serving as programmable connectors between disaggregated computing resources, Conspirator facilitates efficient data transfer without the involvement of host CPUs and hence circumvents the potential bottlenecks there. Conspirator further integrates a novel scheduling algorithm that takes into consideration of the heterogeneity of accelerators and adapts to changing workload dynamics, enabling the flexibility to mitigate the second bottleneck. Our evaluation demonstrates that Conspirator may provide a 15% end-to-end completion time reduction compared to RDMA-based alternatives while being 17% more cost-effective and 44% more power-efficient. Our proposed scheduler also helps to save 33% GPU hours compared to naive GPU-sharing schedulers by making closeto-optimal decisions while taking much less time than the optimal NP-Hard scheduler.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- HiDPU: A DPU-Oriented Hybrid Indexing Scheme for Disaggregated Storage SystemsWenbin Zhu, Zhaoyan Shen, Qian Wei, Renhai Chen 等FAST 2025 · 被引用 11 次
- DDoS Detection at the Scale of One Hundred TbpsYunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang 等NSDI 2026 · 被引用 2 次
- SG-IOV: Socket-Granular I/O Virtualization for SmartNIC-Based Container NetworksChenxingyu Zhao, Hongtao Zhang, Jaehong Min, Shengkai Lin 等ASPLOS 2026 · 被引用 2 次
- HybridMesh: A Hardware-software Hybrid Approach for Accelerating Service Mesh IngressMyoungsung You, Jaehyun Nam, Minjae Seo, Taejune Park 等NSDI 2026 · 被引用 2 次
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min 等SIGCOMM 2026 · 被引用 1 次
它引用的顶会 Paper20
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 被引用 325 次
- AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload RebalancingTong Geng, Ang Li, Runbin Shi, Chunshu Wu 等MICRO 2020 · 被引用 299 次
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee 等OSDI 2020 · 被引用 286 次
相关 Paper
- Tai Chi: A General High-Efficiency Scheduling Framework for SmartNICs in Hyperscale CloudsBang Di, Yun Xu, Kaijie Guo, Yibin Shen 等SOSP 2025
- AlNiCo: SmartNIC-accelerated Contention-aware Request Scheduling for Transaction ProcessingJunru Li, Youyou Lu, Qing Wang, Jiazhen Lin 等USENIX ATC 2022 · 被引用 17 次
- FlexDriver: a network driver for your acceleratorHaggai Eran, Maxim Fudim, Gabi Malka, Gal Shalom 等ASPLOS 2022 · 被引用 14 次
- Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUsDimitrios Giouroukis, Dwi P. A. Nugroho, Varun Pandey, Steffen Zeuch 等VLDB 2025 · 被引用 3 次
- Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AIMikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman 等SC 2024 · 被引用 15 次
