A Distributional Framework For Data Valuation
Amirata Ghorbani, Michael P. Kim, James Zou
Abstract
Shapley value is a classic notion from game theory, historically used to quantify the contributions of individuals within groups, and more recently applied to assign values to data points when training machine learning models. Despite its foundational role, a key limitation of the data Shapley framework is that it only provides valuations for points within a fixed data set. It does not account for statistical aspects of the data and does not give a way to reason about points outside the data set. To address these limitations, we propose a novel framework -- distributional Shapley -- where the value of a point is defined in the context of an underlying data distribution. We prove that distributional Shapley has several desirable statistical properties; for example, the values are stable under perturbations to the data points themselves and to the underlying data distribution. We leverage these properties to develop a new algorithm for estimating values from data, which comes with formal guarantees and runs two orders of magnitude faster than state-of-the-art algorithms for computing the (non-distributional) data Shapley values. We apply distributional Shapley to diverse data sets and demonstrate its utility in a data market setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71b8f278-68be-46cb-aaeb-0231936d4d3dCited by top-tier papers54
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 158 citations
- DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang LowICML 2022 · 71 citations
- Measuring the Effect of Training Data on Deep Learning Predictions via Randomized ExperimentsJinkun Lin, Anqi Zhang, Mathias Lécuyer, Jinyang Li et al.ICML 2022 · 70 citations
- Improving Fairness for Data Valuation in Horizontal Federated LearningZhenan Fan, Huang Fang, Zirui Zhou, Jian Pei et al.ICDE 2022 · 68 citations
- WeightedSHAP: analyzing and improving Shapley based feature attributionsYongchan Kwon, James Y. ZouNeurIPS 2022 · 60 citations
Builds on1
Related papers
- Localized Data Shapley: Accelerating Valuation for Nearest Neighbor AlgorithmsGuangyi Zhang, Yanhao Wang, Chengliang Chai, Qiyu Liu et al.NeurIPS 2025 · 1 citation
- DU-Shapley: A Shapley Value Proxy for Efficient Dataset ValuationFelipe Garrido-Lucero, Benjamin Heymann, Maxime Vono, Patrick Loiseau et al.NeurIPS 2024 · 19 citations
- Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling AlgorithmBoxin Zhao, Boxiang Lyu, Raul Castro Fernandez, Mladen KolarICML 2023 · 14 citations
- EcoVal: An Efficient Data Valuation Framework for Machine LearningAyush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan et al.KDD 2024 · 3 citations
- Dynamic Shapley Value ComputationJiayao Zhang, Haocheng Xia, Qiheng Sun, Jinfei Liu et al.ICDE 2023 · 20 citations
