HyperPrism: An Adaptive Non-linear Aggregation Framework for Distributed Machine Learning over Non-IID Data and Time-varying Communication Links
Haizhou Du, Yijian Chen, Ryan Yang, Yuchen Li, Linghe Kong
摘要
While Distributed Machine Learning (DML) has been widely used to achieve decent performance, it is still challenging to take full advantage of data and devices distributed at multiple vantage points to adapt and learn; this is because the current linear aggregation paradigm cannot solve inter-model divergence caused by (1) heterogeneous learning data at different devices ( i.e. , non-IID data) and (2) in the case of time-varying communication links, the limited ability for devices to reconcile model divergence. In this paper, we present a non-linear class aggregation framework HyperPrism that leverages Kolmogorov Means to conduct distributed mirror descent with the averaging occurring within the mirror descent dual space; HyperPrism selects the degree for a Weighted Power Mean (WPM), a subset of the Kolmogorov Means, each round. Moreover, HyperPrism can adaptively choose different mapping for different layers of the local model with a dedicated hyper-network per device, achieving automatic optimization of DML in high divergence settings. We perform rigorous analysis and experimental evaluations to demonstrate the effectiveness of adaptive, mirror-mapping DML. In particular, we extend the generalizability of existing related works and position them as special cases within HyperPrism. For practitioners, the strength of HyperPrism is in making feasible the possibility of distributed asynchronous training with minimal communication. Our experimental results show HyperPrism can improve the convergence speed up to 98.63% and scale well to more devices compared with the state-of-the-art, all with little additional computation overhead compared to traditional linear aggregation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang 等ICLR 2020 · 被引用 2,930 次
- FedProto: Federated Prototype Learning across Heterogeneous ClientsYue Tan, Guodong Long, Lu Liu, Tianyi Zhou 等AAAI 2022 · 被引用 851 次
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi 等ICML 2020 · 被引用 623 次
- Personalized Federated Learning using HypernetworksAviv Shamsian, Aviv Navon, Ethan Fetaya, Gal ChechikICML 2021 · 被引用 452 次
- Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated LearningLin Zhang, Li Shen, Liang Ding, Dacheng Tao 等CVPR 2022 · 被引用 339 次
相关 Paper
- Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary DelaysChang-Wei Shi, Shi-Shang Wang, Wu-Jun LiAAAI 2026
- Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous MLTehila Dahan, Kfir Y. LevyNeurIPS 2024 · 被引用 4 次
- Edge-consensus Learning: Deep Learning on P2P Networks with Nonhomogeneous DataKenta Niwa, Noboru Harada, Guoqiang Zhang, W. Bastiaan KleijnKDD 2020 · 被引用 34 次
- HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model TrainingGeon-Woo Kim, Junbo Li, Shashidhar Gandham, Omar Baldonado 等ICML 2025
- AdaGK-SGD: Adaptive Global Knowledge Guided Distributed Stochastic Gradient DescentHangyu Ye, Weiying Xie, Yunsong Li, Leyuan FangAAAI 2025 · 被引用 1 次
