Efficient Data Valuation Approximation in Federated Learning: A Sampling-Based Approach
Shuyue Wei, Yongxin Tong, Zimu Zhou, Tianran He, Yi Xu
Abstract
Federated learning (FL) has emerged as a prominent distributed learning paradigm to utilize datasets across multiple data providers. In FL, cross-silo data providers often hesitate to share their high-quality dataset unless their data value can be fairly assessed. Shapley value (SV) has been advocated as the standard metric for data valuation in FL due to its desirable properties. However, the computational overhead of SV is prohibitive in practice, as it inherently requires training and evaluating an FL model across an exponential number of dataset combinations. Furthermore, existing solutions fail to achieve high accuracy and efficiency, making practical use of SV still out of reach, because they ignore choosing suitable computation scheme for approximation framework and overlook the property of utility function in FL. We first propose a unified stratified-sampling framework for two widely-used schemes. Then, we analyze and choose the more promising scheme under the FL linear regression assumption. After that, we identify a phenomenon termed key combinations, where only limited dataset combinations have a high-impact on final data value. Building on these insights, we propose a practical approximation algorithm, IPSS, which strategically selects high-impact dataset combinations rather than evaluating all possible combinations, thus substantially reducing time cost with minor approximation error. Furthermore, we conduct extensive evaluations on the FL benchmark datasets to demonstrate that our proposed algorithm outperforms a series of representative baselines in terms of efficiency and effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96a2eab7-3314-43b3-bc09-21c3b89bfd4eBuilds on18
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi et al.ICML 2020 · 3,875 citations
- Federated Learning on Non-IID Data Silos: An Experimental StudyQinbin Li, Yiqun Diao, Quan Chen, Bingsheng HeICDE 2022 · 1,110 citations
- Personalized Cross-Silo Federated Learning on Non-IID DataYutao Huang, Lingyang Chu, Zirui Zhou, Lanjun Wang et al.AAAI 2021 · 816 citations
- Privacy Preserving Vertical Federated Learning for Tree-based ModelsYuncheng Wu, Shaofeng Cai, Xiaokui Xiao, Gang Chen et al.VLDB 2020 · 259 citations
- Practical Federated Gradient Boosting Decision TreesQinbin Li, Zeyi Wen, Bingsheng HeAAAI 2020 · 215 citations
Related papers
- Efficient Participant Contribution Evaluation for Horizontal and Vertical Federated LearningJunhao Wang, Lan Zhang, Anran Li, Xuanke You et al.ICDE 2022 · 40 citations
- Fair and Efficient Contribution Valuation for Vertical Federated LearningZhenan Fan, Huang Fang, Xinglu Wang, Zirui Zhou et al.ICLR 2024 · 33 citations
- FairFed: Improving Fairness and Efficiency of Contribution Evaluation in Federated Learning via Cooperative Shapley ValueYiqi Liu, Shan Chang, Ye Liu, Bo Li et al.INFOCOM 2024 · 20 citations
- Improving Fairness for Data Valuation in Horizontal Federated LearningZhenan Fan, Huang Fang, Zirui Zhou, Jian Pei et al.ICDE 2022 · 68 citations
- ShapleyFL: Robust Federated Learning Based on Shapley ValueQiheng Sun, Xiang Li, Jiayao Zhang, Li Xiong et al.KDD 2023 · 57 citations
