Hounding Data Diversity: Towards Participant Selection in Vertical Federated Learning
Xiaokai Zhou, Xiao Yan, Fangcheng Fu, Xinyan Li, Hao Huang, Quanqing Xu, Chuanhui Yang, Bo Du, Tieyun Qian, Jiawei Jiang
Abstract
Due to the rising concerns on privacy protection, how to build machine learning models from distributed databases with privacy guarantees has gained more popularity. Vertical federated learning (VFL) trains machine learning models in a privacy-preserving way when the data features are scattered over distributed databases. We study the participant selection problem (PSP) for VFL, which chooses a given number of participants to conduct training while maximizing model accuracy. Compared to training with all participants, PSP can filter out hitch-riders that contribute marginally to model quality and reduce training time by involving fewer participants. To achieve good model accuracy, we formulate PSP as choosing a set of participants that maximizes the likelihood of the data samples. Then, utilizing the k-nearest neighbors (KNN) classifier as the proxy model, we express the likelihood as a function of the selected participants and prove that the function is sub modular. The submodular property is favorable as it can account for the feature diversity among the participants and allows to greedily select the participant with the maximum gain in each step. However, the selection process requires finding the top-k neighbors of a data sample as the basic operation, which is expensive in VFL setting as it involves encrypted communication. As such, we adapt the Fagin's algorithm, a famous top-k query algorithm, to reduce the amount of encrypted communication. We deploy our solution VFPS-SM across five distributed nodes and conduct experiments with 10 datasets and 3 models to evaluate its performance. The results show that VFPS-SM can reduce the end-to-end running time by up to, selection timeand improve model accuracy by 6.0% compared with state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3187341-387b-45de-a0b3-b05714701a86Cited by top-tier papers1
Ask how each one uses itBuilds on32
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen et al.VLDB 2020 · 259 citations
- Feature Inference Attack on Model Predictions in Vertical Federated LearningXinjian Luo, Yuncheng Wu, Xiaokui Xiao, Beng Chin OoiICDE 2021 · 212 citations
Related papers
- VF-PS: How to Select Important Participants in Vertical Federated Learning, Efficiently and Securely?Jiawei Jiang, Lukas Burkhalter, Fangcheng Fu, Bolin Ding et al.NeurIPS 2022 · 43 citations
- PS-MI: Accurate, Efficient, and Private Data Valuation in Vertical Federated LearningXiaokai Zhou, Xiao Yan, Fangcheng Fu, Ziwen Fu et al.VLDB 2025
- VF2Boost: Very Fast Vertical Federated Gradient Boosting for Cross-Enterprise LearningFangcheng Fu, Yingxia Shao, Lele Yu, Jiawei Jiang et al.SIGMOD 2021 · 69 citations
- LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated LearningTimothy Castiglia, Yi Zhou, Shiqiang Wang, Swanand Kadhe et al.ICML 2023 · 33 citations
- Coresets for Vertical Federated Learning: Regularized Linear Regression and -Means ClusteringLingxiao Huang, Zhize Li, Jialin Sun, Haoyu ZhaoNeurIPS 2022 · 31 citations
