Dynamic Parameter Allocation in Parameter Servers
Alexander Renz-Wieland, Rainer Gemulla, Steffen Zeuch, Volker Markl
Abstract
To keep up with increasing dataset sizes and model complexity, distributed training has become a necessity for large machine learning tasks. Parameter servers ease the implementation of distributed parameter management---a key concern in distributed training---, but can induce severe communication overhead. To reduce communication overhead, distributed machine learning algorithms use techniques to increase parameter access locality (PAL), achieving up to linear speed-ups. We found that existing parameter servers provide only limited support for PAL techniques, however, and therefore prevent efficient training. In this paper, we explore whether and to what extent PAL techniques can be supported, and whether such support is beneficial. We propose to integrate dynamic parameter allocation into parameter servers, describe an efficient implementation of such a parameter server called Lapse, and experimentally compare its performance to existing parameter servers across a number of machine learning tasks. We found that Lapse provides near-linear scaling and can be orders of magnitude faster than existing parameter servers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a56f6c0c-673f-467d-a31f-d02ab488b4faCited by top-tier papers6
- Towards Demystifying Serverless Machine Learning TrainingJiawei Jiang, Shaoduo Gan, Yue Liu, Fanlin Wang et al.SIGMOD 2021 · 107 citations
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak et al.VLDB 2021 · 35 citations
- Parallel Training of Knowledge Graph Embedding Models: A Comparison of TechniquesAdrian Kochsiek, Rainer GemullaVLDB 2022 · 33 citations
- HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model TrainingXupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang et al.SIGMOD 2022 · 24 citations
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter AccessAlexander Renz-Wieland, Rainer Gemulla, Zoi Kaoudi, Volker MarklSIGMOD 2022 · 19 citations
Builds on1
Related papers
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 67 citations
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian et al.NSDI 2024 · 16 citations
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu et al.INFOCOM 2024 · 2 citations
- Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DMLJinkun Geng, Dan Li, Shuai WangICDE 2020 · 5 citations
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai et al.SC 2020 · 11 citations
