Conspirator: SmartNIC-Aided Control Plane for Distributed ML Workloads
Yunming Xiao, Diman Zad Tootaghaj, Aditya Dhakal, Lianjie Cao, Puneet Sharma, Aleksandar Kuzmanovic
Abstract
Modern machine learning (ML) workloads heavily depend on distributing tasks across clusters of server CPUs and specialized accelerators, such as GPUs and TPUs, to achieve optimal performance. Nonetheless, prior research has highlighted the inefficient utilization of computing resources in distributed ML, leading to suboptimal performance. This inefficiency primarily stems from CPU bottlenecks and suboptimal accelerator scheduling. Although numerous proposals have been put forward to address these issues individually, none have effectively tackled both inefficiencies simultaneously. In this paper, we introduce Conspirator, an innovative control plane design aimed at alleviating both bottlenecks by harnessing the enhanced computing capabilities of SmartNICs. Following the evolving role of SmartNICs, which have transitioned from their initial function of standard networking task offloading to serving as programmable connectors between disaggregated computing resources, Conspirator facilitates efficient data transfer without the involvement of host CPUs and hence circumvents the potential bottlenecks there. Conspirator further integrates a novel scheduling algorithm that takes into consideration of the heterogeneity of accelerators and adapts to changing workload dynamics, enabling the flexibility to mitigate the second bottleneck. Our evaluation demonstrates that Conspirator may provide a 15% end-to-end completion time reduction compared to RDMA-based alternatives while being 17% more cost-effective and 44% more power-efficient. Our proposed scheduler also helps to save 33% GPU hours compared to naive GPU-sharing schedulers by making closeto-optimal decisions while taking much less time than the optimal NP-Hard scheduler.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e0a6a8c-550f-4735-bf6f-798b0f998c7eCited by top-tier papers7
- HiDPU: A DPU-Oriented Hybrid Indexing Scheme for Disaggregated Storage SystemsWenbin Zhu, Zhaoyan Shen, Qian Wei, Renhai Chen et al.FAST 2025 · 11 citations
- DDoS Detection at the Scale of One Hundred TbpsYunming Xiao, Xijun Luo, Youliang Jiang, Aike Wang et al.NSDI 2026 · 2 citations
- SG-IOV: Socket-Granular I/O Virtualization for SmartNIC-Based Container NetworksChenxingyu Zhao, Hongtao Zhang, Jaehong Min, Shengkai Lin et al.ASPLOS 2026 · 2 citations
- HybridMesh: A Hardware-software Hybrid Approach for Accelerating Service Mesh IngressMyoungsung You, Jaehyun Nam, Minjae Seo, Taejune Park et al.NSDI 2026 · 2 citations
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min et al.SIGCOMM 2026 · 1 citation
Builds on20
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- INFaaS: Automated Model-less Inference ServingFrancisco Romero, Qian Li, Neeraja J. Yadwadkar, Christos KozyrakisUSENIX ATC 2021 · 325 citations
- AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload RebalancingTong Geng, Ang Li, Runbin Shi, Chunshu Wu et al.MICRO 2020 · 299 citations
- Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning WorkloadsDeepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee et al.OSDI 2020 · 286 citations
Related papers
- Tai Chi: A General High-Efficiency Scheduling Framework for SmartNICs in Hyperscale CloudsBang Di, Yun Xu, Kaijie Guo, Yibin Shen et al.SOSP 2025
- AlNiCo: SmartNIC-accelerated Contention-aware Request Scheduling for Transaction ProcessingJunru Li, Youyou Lu, Qing Wang, Jiazhen Lin et al.USENIX ATC 2022 · 17 citations
- FlexDriver: a network driver for your acceleratorHaggai Eran, Maxim Fudim, Gabi Malka, Gal Shalom et al.ASPLOS 2022 · 14 citations
- Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUsDimitrios Giouroukis, Dwi P. A. Nugroho, Varun Pandey, Steffen Zeuch et al.VLDB 2025 · 3 citations
- Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AIMikhail Khalilov, Salvatore Di Girolamo, Marcin Chrapek, Rami Nudelman et al.SC 2024 · 15 citations
