Hounding Data Diversity: Towards Participant Selection in Vertical Federated Learning
Xiaokai Zhou, Xiao Yan, Fangcheng Fu, Xinyan Li, Hao Huang, Quanqing Xu, Chuanhui Yang, Bo Du, Tieyun Qian, Jiawei Jiang
摘要
Due to the rising concerns on privacy protection, how to build machine learning models from distributed databases with privacy guarantees has gained more popularity. Vertical federated learning (VFL) trains machine learning models in a privacy-preserving way when the data features are scattered over distributed databases. We study the participant selection problem (PSP) for VFL, which chooses a given number of participants to conduct training while maximizing model accuracy. Compared to training with all participants, PSP can filter out hitch-riders that contribute marginally to model quality and reduce training time by involving fewer participants. To achieve good model accuracy, we formulate PSP as choosing a set of participants that maximizes the likelihood of the data samples. Then, utilizing the k-nearest neighbors (KNN) classifier as the proxy model, we express the likelihood as a function of the selected participants and prove that the function is sub modular. The submodular property is favorable as it can account for the feature diversity among the participants and allows to greedily select the participant with the maximum gain in each step. However, the selection process requires finding the top-k neighbors of a data sample as the basic operation, which is expensive in VFL setting as it involves encrypted communication. As such, we adapt the Fagin's algorithm, a famous top-k query algorithm, to reduce the amount of encrypted communication. We deploy our solution VFPS-SM across five distributed nodes and conduct experiments with 10 datasets and 3 models to evaluate its performance. The results show that VFPS-SM can reduce the end-to-end running time by up to, selection timeand improve model accuracy by 6.0% compared with state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper32
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 被引用 1,110 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen 等VLDB 2020 · 被引用 259 次
- Feature Inference Attack on Model Predictions in Vertical Federated LearningXinjian Luo, Yuncheng Wu, Xiaokui Xiao, Beng Chin OoiICDE 2021 · 被引用 212 次
相关 Paper
- VF-PS: How to Select Important Participants in Vertical Federated Learning, Efficiently and Securely?Jiawei Jiang, Lukas Burkhalter, Fangcheng Fu, Bolin Ding 等NeurIPS 2022 · 被引用 43 次
- PS-MI: Accurate, Efficient, and Private Data Valuation in Vertical Federated LearningXiaokai Zhou, Xiao Yan, Fangcheng Fu, Ziwen Fu 等VLDB 2025
- VF2Boost: Very Fast Vertical Federated Gradient Boosting for Cross-Enterprise LearningFangcheng Fu, Yingxia Shao, Lele Yu, Jiawei Jiang 等SIGMOD 2021 · 被引用 69 次
- LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated LearningTimothy Castiglia, Yi Zhou, Shiqiang Wang, Swanand Kadhe 等ICML 2023 · 被引用 33 次
- Coresets for Vertical Federated Learning: Regularized Linear Regression and -Means ClusteringLingxiao Huang, Zhize Li, Jialin Sun, Haoyu ZhaoNeurIPS 2022 · 被引用 31 次
