Lune

NSDI2025Top-tier venue

OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud

Ertza Warraich, Omer Shabtai, Khalid Manaa, Shay Vargaftik, Yonatan Piasetzky, Matty Kadosh, Lalith Suresh, Muhammad Shahbaz

2025Year

Abstract

We present OPTIREDUCE, a new collective-communication system for the cloud with bounded, predictable completion times for deep-learning jobs in the presence of varying computation (stragglers) and communication (congestion and gradient drops) variabilities. OPTIREDUCE exploits the inherent resiliency and the stochastic nature of distributed deep-learning (DDL) training and fine-tuning to work with approximated (or lost) gradients-providing an efficient balance between tail performance and the resulting accuracy of the trained models.

Exploiting this domain-specific characteristic of DDL, OPTIREDUCE introduces (1) mechanisms (e.g., unreliable bounded transport with adaptive timeout) to improve the DDL jobs' tail execution time, and (2) strategies (e.g., Transpose AllReduce and Hadamard Transform) to mitigate the impact of dropped gradient entries on model accuracy. Our evaluation shows that OPTIREDUCE achieves 70% and 30% faster timeto-accuracy (TTA), on average, when operating in shared, cloud environments (e.g., CloudLab) compared to Gloo and NCCL, respectively.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3889af3f-78e1-470f-8219-00d8b9abd4e7

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines