HyperPrism: An Adaptive Non-linear Aggregation Framework for Distributed Machine Learning over Non-IID Data and Time-varying Communication Links
Haizhou Du, Yijian Chen, Ryan Yang, Yuchen Li, Linghe Kong
Abstract
While Distributed Machine Learning (DML) has been widely used to achieve decent performance, it is still challenging to take full advantage of data and devices distributed at multiple vantage points to adapt and learn; this is because the current linear aggregation paradigm cannot solve inter-model divergence caused by (1) heterogeneous learning data at different devices ( i.e. , non-IID data) and (2) in the case of time-varying communication links, the limited ability for devices to reconcile model divergence. In this paper, we present a non-linear class aggregation framework HyperPrism that leverages Kolmogorov Means to conduct distributed mirror descent with the averaging occurring within the mirror descent dual space; HyperPrism selects the degree for a Weighted Power Mean (WPM), a subset of the Kolmogorov Means, each round. Moreover, HyperPrism can adaptively choose different mapping for different layers of the local model with a dedicated hyper-network per device, achieving automatic optimization of DML in high divergence settings. We perform rigorous analysis and experimental evaluations to demonstrate the effectiveness of adaptive, mirror-mapping DML. In particular, we extend the generalizability of existing related works and position them as special cases within HyperPrism. For practitioners, the strength of HyperPrism is in making feasible the possibility of distributed asynchronous training with minimal communication. Our experimental results show HyperPrism can improve the convergence speed up to 98.63% and scale well to more devices compared with the state-of-the-art, all with little additional computation overhead compared to traditional linear aggregation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext faba7db3-113f-40cc-b702-d480527b2dc5Builds on17
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- FedProto: Federated Prototype Learning across Heterogeneous ClientsYue Tan, Guodong Long, Lu Liu, Tianyi Zhou et al.AAAI 2022 · 851 citations
- A Unified Theory of Decentralized SGD with Changing Topology and Local UpdatesAnastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi et al.ICML 2020 · 623 citations
- Personalized Federated Learning using HypernetworksAviv Shamsian, Aviv Navon, Ethan Fetaya, Gal ChechikICML 2021 · 452 citations
- Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated LearningLin Zhang, Li Shen, Liang Ding, Dacheng Tao et al.CVPR 2022 · 339 citations
Related papers
- Ordered Local Momentum for Asynchronous Distributed Learning Under Arbitrary DelaysChang-Wei Shi, Shi-Shang Wang, Wu-Jun LiAAAI 2026
- Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous MLTehila Dahan, Kfir Y. LevyNeurIPS 2024 · 4 citations
- Edge-consensus Learning: Deep Learning on P2P Networks with Nonhomogeneous DataKenta Niwa, Noboru Harada, Guoqiang Zhang, W. Bastiaan KleijnKDD 2020 · 34 citations
- HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model TrainingGeon-Woo Kim, Junbo Li, Shashidhar Gandham, Omar Baldonado et al.ICML 2025
- AdaGK-SGD: Adaptive Global Knowledge Guided Distributed Stochastic Gradient DescentHangyu Ye, Weiying Xie, Yunsong Li, Leyuan FangAAAI 2025 · 1 citation
