PS-MI: Accurate, Efficient, and Private Data Valuation in Vertical Federated Learning
Xiaokai Zhou, Xiao Yan, Fangcheng Fu, Ziwen Fu, Tieyun Qian, Yuanyuan Zhu, Qinbo Zhang, Bin Cui, Jiawei Jiang
Abstract
Vertical federated learning (VFL) trains models when multiple databases (a.k.a participants) hold different features of the same set of samples. By quantifying each participant's contribution to model training, data valuation can prevent hitch-riders and reward the instrumental parties. However, vertical federated data valuation (VFDV) is challenging because it needs to be accurate and efficient while protecting participant data privacy. In this paper, we propose a method meeting all three requirements by using projection and sampling for mutual information estimation (thus dubbed PS-MI). In particular, we first show that the utility of a participant set (a.k.a a coalition ) can be expressed as the mutual information (MI) between their features and the target labels. MI is favorable because it does not depend on the model to train (i.e., model-agnostic ) and can be estimated via k -nearest neighbor (KNN). To run KNN, instead of using costly homomorphic encryption to protect data privacy, we apply simple random projection to participant features before distance computation. We prove that random projection ensures differential privacy and preserves unbiased distance estimates. Since the contribution of a participant involves many coalitions, we adopt stratified sampling to reduce the number of coalitions while controlling estimation variance. To further improve efficiency, we incorporate optimizations including using locality sensitive hashing (LSH) to prune kNN candidates, batching kNN candidate checking for multiple coalitions, and adaptive early termination for utility evaluation. We compare PS-MI with 5 state-of-the-art VFDV methods. The results show that PS-MI yields higher accuracy and shorter running time than the baselines, and the maximum speedup can be 592×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d78e2ef1-5aaa-49ea-9770-d715c9d9732bCited by top-tier papers1
Ask how each one uses itBuilds on16
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 2,107 citations
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- Fast Private Set Intersection from Homomorphic EncryptionHao Chen, Kim Laine, Peter RindalCCS 2017 · 446 citations
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen et al.VLDB 2020 · 259 citations
- Feature Inference Attack on Model Predictions in Vertical Federated LearningXinjian Luo, Yuncheng Wu, Xiaokui Xiao, Beng Chin OoiICDE 2021 · 212 citations
Related papers
- Hounding Data Diversity: Towards Participant Selection in Vertical Federated LearningXiaokai Zhou, Xiao Yan, Fangcheng Fu, Xinyan Li et al.ICDE 2025 · 1 citation
- VF-PS: How to Select Important Participants in Vertical Federated Learning, Efficiently and Securely?Jiawei Jiang, Lukas Burkhalter, Fangcheng Fu, Bolin Ding et al.NeurIPS 2022 · 43 citations
- Fair and Efficient Contribution Valuation for Vertical Federated LearningZhenan Fan, Huang Fang, Xinglu Wang, Zirui Zhou et al.ICLR 2024 · 33 citations
- FedMix: Boosting with Data Mixture for Vertical Federated LearningYihang Cheng, Lan Zhang, Junyang Wang, Xiaokai Chu et al.ICDE 2024 · 6 citations
- Efficient Data Valuation Approximation in Federated Learning: A Sampling-Based ApproachShuyue Wei, Yongxin Tong, Zimu Zhou, Tianran He et al.ICDE 2025
