ATA: Adaptive Task Allocation for Efficient Resource Management in Distributed Machine Learning
Arto Maranjyan, El Mehdi Saad, Peter Richtárik, Francesco Orabona
Abstract
Asynchronous methods are fundamental for parallelizing computations in distributed machine learning. They aim to accelerate training by fully utilizing all available resources. However, their greedy approach can lead to inefficiencies using more computation than required, especially when computation times vary across devices. If the computation times were known in advance, training could be fast and resource-efficient by assigning more tasks to faster workers. The challenge lies in achieving this optimal allocation without prior knowledge of the computation time distributions. In this paper, we propose ATA (Adaptive Task Allocation), a method that adapts to heterogeneous and random distributions of worker computation times. Through rigorous theoretical analysis, we show that ATA identifies the optimal task allocation and performs comparably to methods with prior knowledge of computation times. Experimental results further demonstrate that ATA is resource-efficient, significantly reducing costs compared to the greedy approach, which can be arbitrarily expensive depending on the number of workers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92a93ef7-1f5d-431f-abea-ec06c5b428f8Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Ringmaster ASGD: The First Asynchronous SGD with Optimal Time ComplexityArto Maranjyan, Alexander Tyurin, Peter RichtárikICML 2025
- Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated LearningAnastasia Koloskova, Sebastian U. Stich, Martin JaggiNeurIPS 2022 · 131 citations
- Communication-efficient Decentralized Machine Learning over Heterogeneous NetworksPan Zhou, Qian Lin, Dumitrel Loghin, Beng Chin Ooi et al.ICDE 2021 · 73 citations
- Training Job Placement in Clusters with Statistical In-Network AggregationBohan Zhao, Wei Xu, Shuo Liu, Yang Tian et al.ASPLOS 2024 · 17 citations
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun et al.SC 2023 · 16 citations
