Dynamic Parameter Allocation in Parameter Servers
Alexander Renz-Wieland, Rainer Gemulla, Steffen Zeuch, Volker Markl
摘要
To keep up with increasing dataset sizes and model complexity, distributed training has become a necessity for large machine learning tasks. Parameter servers ease the implementation of distributed parameter management---a key concern in distributed training---, but can induce severe communication overhead. To reduce communication overhead, distributed machine learning algorithms use techniques to increase parameter access locality (PAL), achieving up to linear speed-ups. We found that existing parameter servers provide only limited support for PAL techniques, however, and therefore prevent efficient training. In this paper, we explore whether and to what extent PAL techniques can be supported, and whether such support is beneficial. We propose to integrate dynamic parameter allocation into parameter servers, describe an efficient implementation of such a parameter server called Lapse, and experimentally compare its performance to existing parameter servers across a number of machine learning tasks. We found that Lapse provides near-linear scaling and can be orders of magnitude faster than existing parameter servers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Towards Demystifying Serverless Machine Learning TrainingJiawei Jiang, Shaoduo Gan, Yue Liu, Fanlin Wang 等SIGMOD 2021 · 被引用 107 次
- Distributed Deep Learning on Data Systems: A Comparative Analysis of ApproachesYuhao Zhang, Frank Mcquillan, Nandish Jayaram, Nikhil Kak 等VLDB 2021 · 被引用 35 次
- Parallel Training of Knowledge Graph Embedding Models: A Comparison of TechniquesAdrian Kochsiek, Rainer GemullaVLDB 2022 · 被引用 33 次
- HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model TrainingXupeng Miao, Yining Shi, Hailin Zhang, Xin Zhang 等SIGMOD 2022 · 被引用 24 次
- NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter AccessAlexander Renz-Wieland, Rainer Gemulla, Zoi Kaoudi, Volker MarklSIGMOD 2022 · 被引用 19 次
它引用的顶会 Paper1
相关 Paper
- Preemptive All-reduce Scheduling for Expediting Distributed DNN TrainingYixin Bao, Yanghua Peng, Yangrui Chen, Chuan WuINFOCOM 2020 · 被引用 67 次
- Accelerating Neural Recommendation Training with Embedding SchedulingChaoliang Zeng, Xudong Liao, Xiaodian Cheng, Han Tian 等NSDI 2024 · 被引用 16 次
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu 等INFOCOM 2024 · 被引用 2 次
- Fela: Incorporating Flexible Parallelism and Elastic Tuning to Accelerate Large-Scale DMLJinkun Geng, Dan Li, Shuai WangICDE 2020 · 被引用 5 次
- Herring: rethinking the parameter server at scale for the cloudIndu Thangakrishnan, Derya Cavdar, Can Karakus, Piyush Ghai 等SC 2020 · 被引用 11 次
