NuPS: A Parameter Server for Machine Learning with Non-Uniform Parameter Access
Alexander Renz-Wieland, Rainer Gemulla, Zoi Kaoudi, Volker Markl
摘要
Parameter servers (PSs) facilitate the implementation of distributed training for large machine learning tasks. In this paper, we argue that existing PSs are inefficient for tasks that exhibit non-uniform parameter access; their performance may even fall behind that of single node baselines. We identify two major sources of such non-uniform access: skew and sampling. Existing PSs are ill-suited for managing skew because they uniformly apply the same parameter management technique to all parameters. They are inefficient for sampling because the PS is oblivious to the associated randomized accesses and cannot exploit locality. To overcome these performance limitations, we introduce NuPS, a novel PS architecture that (i) integrates multiple management techniques and employs a suitable technique for each parameter and (ii) supports sampling directly via suitable sampling primitives and sampling schemes that allow for a controlled quality-efficiency trade-off. In our experimental study, NuPS outperformed existing PSs by up to one order of magnitude and provided up to linear scalability across multiple machine learning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachZhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu 等SIGMOD 2024 · 被引用 14 次
- PetPS: Supporting Huge Embedding Models with Persistent MemoryMinhui Xie, Youyou Lu, Qing Wang, Yangyang Feng 等VLDB 2023 · 被引用 10 次
- Saturn: An Optimized Data System for Multi-Large-Model Deep Learning WorkloadsKabir Nagrecha, Arun KumarVLDB 2024 · 被引用 8 次
- SparDL: Distributed Deep Learning Training with Efficient Sparse CommunicationMinjun Zhao, Yichen Yin, Yuren Mao, Qing Liu 等ICDE 2024 · 被引用 6 次
- FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop MechanismPeng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li 等VLDB 2026
它引用的顶会 Paper8
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- You CAN Teach an Old Dog New Tricks! On Training Knowledge Graph EmbeddingsDaniel Ruffinelli, Samuel Broscheit, Rainer GemullaICLR 2020 · 被引用 238 次
- Understanding Negative Sampling in Graph Representation LearningZhen Yang, Ming Ding, Chang Zhou, Hongxia Yang 等KDD 2020 · 被引用 172 次
- DGL-KE: Training Knowledge Graph Embeddings at ScaleDa Zheng, Xiang Song, Chao Ma, Zeyuan Tan 等SIGIR 2020 · 被引用 132 次
- Parallel Training of Knowledge Graph Embedding Models: A Comparison of TechniquesAdrian Kochsiek, Rainer GemullaVLDB 2022 · 被引用 33 次
相关 Paper
- Dynamic Parameter Allocation in Parameter ServersAlexander Renz-Wieland, Rainer Gemulla, Steffen Zeuch, Volker MarklVLDB 2020 · 被引用 18 次
- Gsyn: Reducing Staleness and Communication Waiting via Grouping-based Synchronization for Distributed Deep LearningYijun Li, Jiawei Huang, Zhaoyi Li, Jingling Liu 等INFOCOM 2024 · 被引用 2 次
- Distributed Machine Learning through Heterogeneous Edge SystemsHanpeng Hu, Dan Wang, Chuan WuAAAI 2020 · 被引用 48 次
- Learning Efficient Parameter Server Synchronization Policies for Distributed SGDRong Zhu, Sheng Yang, Andreas Pfadler, Zhengping Qian 等ICLR 2020 · 被引用 9 次
- HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model TrainingGeon-Woo Kim, Junbo Li, Shashidhar Gandham, Omar Baldonado 等ICML 2025
