PS-MI: Accurate, Efficient, and Private Data Valuation in Vertical Federated Learning
Xiaokai Zhou, Xiao Yan, Fangcheng Fu, Ziwen Fu, Tieyun Qian, Yuanyuan Zhu, Qinbo Zhang, Bin Cui, Jiawei Jiang
摘要
Vertical federated learning (VFL) trains models when multiple databases (a.k.a participants) hold different features of the same set of samples. By quantifying each participant's contribution to model training, data valuation can prevent hitch-riders and reward the instrumental parties. However, vertical federated data valuation (VFDV) is challenging because it needs to be accurate and efficient while protecting participant data privacy. In this paper, we propose a method meeting all three requirements by using projection and sampling for mutual information estimation (thus dubbed PS-MI). In particular, we first show that the utility of a participant set (a.k.a a coalition ) can be expressed as the mutual information (MI) between their features and the target labels. MI is favorable because it does not depend on the model to train (i.e., model-agnostic ) and can be estimated via k -nearest neighbor (KNN). To run KNN, instead of using costly homomorphic encryption to protect data privacy, we apply simple random projection to participant features before distance computation. We prove that random projection ensures differential privacy and preserves unbiased distance estimates. Since the contribution of a participant involves many coalitions, we adopt stratified sampling to reduce the number of coalitions while controlling estimation variance. To further improve efficiency, we incorporate optimizations including using locality sensitive hashing (LSH) to prune kNN candidates, batching kNN candidate checking for multiple coalitions, and adaptive early termination for utility evaluation. We compare PS-MI with 5 state-of-the-art VFDV methods. The results show that PS-MI yields higher accuracy and shorter running time than the baselines, and the maximum speedup can be 592×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 被引用 2,107 次
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 被引用 476 次
- Fast Private Set Intersection from Homomorphic EncryptionHao Chen, Kim Laine, Peter RindalCCS 2017 · 被引用 446 次
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen 等VLDB 2020 · 被引用 259 次
- Feature Inference Attack on Model Predictions in Vertical Federated LearningXinjian Luo, Yuncheng Wu, Xiaokui Xiao, Beng Chin OoiICDE 2021 · 被引用 212 次
相关 Paper
- Hounding Data Diversity: Towards Participant Selection in Vertical Federated LearningXiaokai Zhou, Xiao Yan, Fangcheng Fu, Xinyan Li 等ICDE 2025 · 被引用 1 次
- VF-PS: How to Select Important Participants in Vertical Federated Learning, Efficiently and Securely?Jiawei Jiang, Lukas Burkhalter, Fangcheng Fu, Bolin Ding 等NeurIPS 2022 · 被引用 43 次
- Fair and Efficient Contribution Valuation for Vertical Federated LearningZhenan Fan, Huang Fang, Xinglu Wang, Zirui Zhou 等ICLR 2024 · 被引用 33 次
- FedMix: Boosting with Data Mixture for Vertical Federated LearningYihang Cheng, Lan Zhang, Junyang Wang, Xiaokai Chu 等ICDE 2024 · 被引用 6 次
- Efficient Data Valuation Approximation in Federated Learning: A Sampling-Based ApproachShuyue Wei, Yongxin Tong, Zimu Zhou, Tianran He 等ICDE 2025
