Hoplite: efficient and fault-tolerant collective communication for task-based distributed systems
Siyuan Zhuang, Zhuohan Li, Danyang Zhuo, Stephanie Wang, Eric Liang, Robert Nishihara, Philipp Moritz, Ion Stoica
Abstract
Task-based distributed frameworks (e.g., Ray, Dask, Hydro) have become increasingly popular for distributed applications that contain asynchronous and dynamic workloads, including asynchronous gradient descent, reinforcement learning, and model serving. As more data-intensive applications move to run on top of task-based systems, collective communication efficiency has become an important problem. Unfortunately, traditional collective communication libraries (e.g., MPI, Horovod, NCCL) are an ill fit, because they require the communication schedule to be known before runtime and they do not provide fault tolerance.
We design and implement Hoplite, an efficient and fault-tolerant collective communication layer for task-based distributed systems.
Our key technique is to compute data transfer schedules on the fly and execute the schedules efficiently through fine-grained pipelining. At the same time, when a task fails, the data transfer schedule adapts quickly to allow other tasks to keep making progress. We apply Hoplite to a popular task-based distributed framework, Ray. We show that Hoplite speeds up asynchronous stochastic gradient descent, reinforcement learning, and serving an ensemble of machine learning models that are difficult to execute efficiently with traditional collective communication by up to 7.8x, 3.9x, and 3.3x, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32a284d7-51a5-4d71-b9c7-1564dec80efeCited by top-tier papers2
- Exoshuffle: An Extensible Shuffle ArchitectureFrank Sifei Luan, Stephanie Wang, Samyukta Yagati, Sean Kim et al.SIGCOMM 2023 · 4 citations
- TensorHub: Scalable and Elastic Weight Transfer for LLM RL TrainingChenhao Ye, Huaizheng Zhang, Mingcong Han, Baoquan Zhong et al.SOSP 2026
Builds on1
Related papers
- Handling Network Faults in Distributed AI Training: Failover is Now an OptionXin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui et al.EuroSys 2026 · 1 citation
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai et al.SC 2020 · 11 citations
- OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the CloudErtza Warraich, Omer Shabtai, Khalid Manaa, Shay Vargaftik et al.NSDI 2025
- PReCCL: Performant and Resilient Collective Communication via Integrated Inband Telemetry and Workload ReallocationZhiyong Chen, Kaihui Gao, Li Chen, Rui Yan et al.SIGCOMM 2026
- Composing Distributed Computations Through Task and Kernel FusionRohan Yadav, Shiv Sundram, Wonchan Lee, Michael Garland et al.ASPLOS 2025
