FLea: Addressing Data Scarcity and Label Skew in Federated Learning via Privacy-preserving Feature Augmentation
Tong Xia, Abhirup Ghosh, Xinchi Qiu, Cecilia Mascolo
Abstract
Federated Learning (FL) enables model development by leveraging data distributed across numerous edge devices without transferring local data to a central server. However, existing FL methods still face challenges when dealing with scarce and label-skewed data across devices, resulting in local model overfitting and drift, consequently hindering the performance of the global model. In response to these challenges, we propose a pioneering framework called FLea, incorporating the following key components: i) A global feature buffer that stores activation-target pairs shared from multiple clients to support local training. This design mitigates local model drift caused by the absence of certain classes; ii) A feature augmentation approach based on local and global activation mix-ups for local training. This strategy enlarges the training samples, thereby reducing the risk of local overfitting; iii) An obfuscation method to minimize the correlation between intermediate activations and the source data, enhancing the privacy of shared features. To verify the superiority of FLea, we conduct extensive experiments using a wide range of data modalities, simulating different levels of local data scarcity and label skew. The results demonstrate that FLea consistently outperforms state-of-the-art FL counterparts (among 13 of the experimented 18 settings, the improvement is over 5%) while concurrently mitigating the privacy vulnerabilities associated with shared features. Code is available at https://github.com/XTxiatong/FLea.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5508369e-5d0d-4a8d-933b-9cdd64f11e42Builds on17
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
- FedProto: Federated Prototype Learning across Heterogeneous ClientsYue Tan, Guodong Long, Lu Liu, Tianyi Zhou et al.AAAI 2022 · 851 citations
- No Fear of Heterogeneity: Classifier Calibration for Federated Learning with Non-IID DataMi Luo, Fei Chen, Dapeng Hu, Yifan Zhang et al.NeurIPS 2021 · 510 citations
Related papers
- FedMix: Approximation of Mixup under Mean Augmented Federated LearningTehrim Yoon, Sumin Shin, Sung Ju Hwang, Eunho YangICLR 2021 · 226 citations
- DapperFL: Domain Adaptive Federated Learning with Model Fusion Pruning for Edge DevicesYongzhe Jia, Xuyun Zhang, Hongsheng Hu, Kim-Kwang Raymond Choo et al.NeurIPS 2024 · 14 citations
- FedCE: Personalized Federated Learning Method based on Clustering EnsemblesLuxin Cai, Naiyue Chen, Yuanzhouhan Cao, Jiahuan He et al.ACM MM 2023 · 27 citations
- FedASMU: Efficient Asynchronous Federated Learning with Dynamic Staleness-Aware Model UpdateJi Liu, Juncheng Jia, Tianshi Che, Chao Huo et al.AAAI 2024 · 87 citations
- Enhancing Clustered Federated Learning: Integration of Strategies and Improved MethodologiesYongxin Guo, Xiaoying Tang, Tao LinICLR 2025
