FedUSD: Unbiased Synthetic Data for Federated Learning
Weiying Xie, Chenhe Hao, Haozhi Shi, Jitao Ma, Daixun Li, Jiazhe Li, Hengyi Wang, Leyuan Fang, Yunsong Li
Abstract
Aggregation-Free Federated Learning enables joint training by sharing synthetic data, aiming to eliminate data heterogeneity across clients. However, existing methods fail to explicitly separate the principal and residual components of dataset, leading to biased synthetic data. In this paper, we propose a novel Unbiased Synthetic Data optimization method FedUSD for Aggregation-Free Federated Learning, which is achieved by exploring the High-energy Orthogonal Base (HOB) and variance of dataset in feature space. Our FedUSD is inspired by the discovery that principal component concentrates in HOB while residual component independently reflects in variance, regardless of networks. Based on the observation, we develop a method that mathematically optimizes synthetic data by matching both HOB and variance with those of real data. Besides, we experimentally show the superior effectiveness of leveraging HOB and variance to separately extract the principal and residual components over existing methods. We also theoretically prove that FedUSD achieves unbiased synthetic data and thus convergence. Without introducing any constraints, FedUSD thereby yields significant improvements over the state-of-the-arts in terms of global model performance, under equivalent communicational costs. For example, on the SVHN dataset, FedUSD improves 6.74% to 30.82% which is higher than others with Dirichlet coefficient .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a676c96a-82a4-441b-b609-43f4300bed2fBuilds on27
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated OptimizationJianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi et al.NeurIPS 2020 · 2,231 citations
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 1,615 citations
Related papers
- Exploiting Label Skews in Federated Learning with Model ConcatenationYiqun Diao, Qinbin Li, Bingsheng HeAAAI 2024 · 39 citations
- Fake It Till Make It: Federated Learning with Consensus-Oriented GenerationRui Ye, Yaxin Du, Zhenyang Ni, Yanfeng Wang et al.ICLR 2024 · 11 citations
- Bridging Generalization Gap of Heterogeneous Federated Clients Using Generative ModelsZiru Niu, Hai Dong, A. K. QinICLR 2026 · 3 citations
- FedSMU: Communication-Efficient and Generalization-Enhanced Federated Learning through Symbolic Model UpdatesXinyi Lu, Hao Zhang, Chenglin Li, Weijia Lu et al.ICML 2025
- Covariances for Free: Exploiting Mean Distributions for Training-free Federated LearningDipam Goswami, Simone Magistri, Kai Wang, Bartlomiej Twardowski et al.NeurIPS 2025 · 3 citations
