Data-Free Black-Box Federated Learning via Zeroth-Order Gradient Estimation
Xinge Ma, Jin Wang, Xuejie Zhang
Abstract
Federated learning (FL) enables decentralized clients to collaboratively train a global model under the orchestration of a central server without exposing their individual data. However, the iterative exchange of model parameters between the server and clients imposes heavy communication burdens, risks potential privacy leakage, and even precludes collaboration among heterogeneous clients. Distillation-based FL tackles these challenges by exchanging low-dimensional model outputs rather than model parameters, yet it highly relies on a task-relevant auxiliary dataset that is often not available in practice. Data-free FL attempts to overcome this limitation by training a server-side generator to directly synthesize task-specific data samples for knowledge transfer. However, the update rule of the generator requires clients to share on-device models for white-box access, which greatly compromises the advantages of distillation-based FL. This motivates us to explore a data-free and black-box FL framework via Zeroth-order Gradient Estimation (FedZGE), which estimates the gradients after flowing through on-device models in a black-box optimization manner to complete the training of the generator in terms of fidelity, transferability, diversity, and equilibrium, without involving any auxiliary data or sharing any model parameters, thus combining the advantages of both distillation-based FL and data-free FL. Experiments on large-scale image classification datasets and network architectures demonstrate the superiority of FedZGE in terms of data heterogeneity, model heterogeneity, communication efficiency, and privacy protection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67aab8cd-a644-495e-8fb2-2256307cbef3Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Data-Free Knowledge Distillation for Heterogeneous Federated LearningZhuangdi Zhu, Junyuan Hong, Jiayu ZhouICML 2021 · 957 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated LearningLin Zhang, Li Shen, Liang Ding, Dacheng Tao et al.CVPR 2022 · 339 citations
- Federated Learning from Pre-Trained Models: A Contrastive Learning ApproachYue Tan, Guodong Long, Jie Ma, Lu Liu et al.NeurIPS 2022 · 316 citations
- FedRolex: Model-Heterogeneous Federated Learning with Rolling Sub-Model ExtractionSamiul Alam, Luyang Liu, Ming Yan, Mi ZhangNeurIPS 2022 · 261 citations
Related papers
- DFRD: Data-Free Robustness Distillation for Heterogeneous Federated LearningKangyang Luo, Shuai Wang, Yexuan Fu, Xiang Li et al.NeurIPS 2023 · 64 citations
- FedFree: Breaking Knowledge-sharing Barriers through Layer-wise Alignment in Heterogeneous Federated LearningHaizhou Du, Yiran Xiang, Yiwen Cai, Xiufeng Liu et al.NeurIPS 2025 · 3 citations
- Provably Near-Optimal Federated Ensemble Distillation with Negligible OverheadWon-Jun Jang, Hyeon-Seo Park, Si-Hyeon LeeICML 2025
- Ensemble Distillation for Robust Model Fusion in Federated LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiNeurIPS 2020 · 1,615 citations
- DENSE: Data-Free One-Shot Federated LearningJie Zhang, Chen Chen, Bo Li, Lingjuan Lyu et al.NeurIPS 2022 · 202 citations
