Improved Modelling of Federated Datasets using Mixtures-of-Dirichlet-Multinomials
Jonathan Scott, Áine Cahill
Abstract
In practice, training using federated learning can be orders of magnitude slower than standard centralized training. This severely limits the amount of experimentation and tuning that can be done, making it challenging to obtain good performance on a given task. Server-side proxy data can be used to run training simulations, for instance for hyperparameter tuning. This can greatly speed up the training pipeline by reducing the number of tuning runs to be performed overall on the true clients. However, it is challenging to ensure that these simulations accurately reflect the dynamics of the real federated training. In particular, the proxy data used for simulations often comes as a single centralized dataset without a partition into distinct clients, and partitioning this data in a naive way can lead to simulations that poorly reflect real federated training. In this paper we address the challenge of how to partition centralized data in a way that reflects the statistical heterogeneity of the true federated clients. We propose a fully federated, theoretically justified, algorithm that efficiently learns the distribution of the true clients and observe improved server-side simulations when using the inferred distribution to create simulated clients from the centralized data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efac45e7-44e1-47a2-a6df-86f06d611baeCited by top-tier papers1
Ask how each one uses itBuilds on7
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning ApproachAlireza Fallah, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2020 · 1,354 citations
- An Efficient Framework for Clustered Federated LearningAvishek Ghosh, Jichan Chung, Dong Yin, Kannan RamchandranNeurIPS 2020 · 1,329 citations
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 671 citations
Related papers
- Where to Begin? On the Impact of Pre-Training and Initialization in Federated LearningJohn Nguyen, Jianyu Wang, Kshitiz Malik, Maziar Sanjabi et al.ICLR 2023 · 13 citations
- Benchmarking Algorithms for Federated Domain GeneralizationRuqi Bai, Saurabh Bagchi, David I. InouyeICLR 2024 · 20 citations
- Tackling System and Statistical Heterogeneity for Federated Learning with Adaptive Client SamplingBing Luo, Wenli Xiao, Shiqiang Wang, Jianwei Huang et al.INFOCOM 2022 · 224 citations
- DYNAFED: Tackling Client Data Heterogeneity with Global DynamicsRenjie Pi, Weizhong Zhang, Yueqi Xie, Jiahui Gao et al.CVPR 2023
- ProxyFL: A Proxy-Guided Framework for Federated Semi-Supervised LearningDuowen Chen, Yan WangCVPR 2026
