Distributionally Robust Data Valuation
Xiaoqiang Lin, Xinyi Xu, Zhaoxuan Wu, See-Kiong Ng, Bryan Kian Hsiang Low
Abstract
Data valuation quantifies the contribution of each data point to the performance of a machine learning model. Existing works typically define the value of data by its improvement of the validation performance of the trained model. However, this approach can be impractical to apply in collaborative machine learning and data marketplace since it is difficult for the parties/buyers to agree on a common validation dataset or determine the exact validation distribution a priori. To address this, we propose a distributionally robust data valuation approach to perform data valuation without known/fixed validation distributions. Our approach defines the value of data by its improvement of the distributionally robust generalization error (DRGE), thus providing a worst-case performance guarantee without a known/fixed validation distribution. However, since computing DRGE directly is infeasible, we propose using model deviation as a proxy for the marginal improvement of DRGE (for kernel regression and neural networks) to compute data values. Furthermore, we identify a notion of uniqueness where low uniqueness characterizes low-value data. We empirically demonstrate that our approach outperforms existing data valuation approaches in data selection and data removal tasks on real-world datasets (e.g., housing price prediction, diabetes hospitalization prediction).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25efe79b-7b83-4052-9a56-cecc2376b77fCited by top-tier papers5
- Helpful or Harmful Data? Fine-tuning-free Shapley Attribution for Explaining Language Model PredictionsJingtan Wang, Xiaoqiang Lin, Rui Qiao, Chuan-Sheng Foo et al.ICML 2024 · 12 citations
- DETAIL: Task DEmonsTration Attribution for Interpretable In-context LearningZijian Zhou, Xiaoqiang Lin, Xinyi Xu, Alok Prakash et al.NeurIPS 2024 · 9 citations
- TETRIS: Optimal Draft Token Selection for Batch Speculative DecodingZhaoxuan Wu, Zijian Zhou, Arun Verma, Alok Prakash et al.ACL 2025 · 7 citations
- Efficient Top-m Data Values Identification for Data SelectionXiaoqiang Lin, Xinyi Xu, See-Kiong Ng, Bryan Kian Hsiang LowICLR 2025
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
Builds on12
- Ditto: Fair and Robust Federated Learning Through PersonalizationTian Li, Shengyuan Hu, Ahmad Beirami, Virginia SmithICML 2021 · 1,313 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 158 citations
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 152 citations
- Dealer: An End-to-End Model Marketplace with Differential PrivacyJinfei Liu, Jian Lou, Junxu Liu, Li Xiong et al.VLDB 2021 · 99 citations
Related papers
- Validation Free and Replication Robust Volume-based Data ValuationXinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang LowNeurIPS 2021 · 89 citations
- LAVA: Data Valuation without Pre-Specified Learning AlgorithmsHoang Anh Just, Feiyang Kang, Tianhao Wang, Yi Zeng et al.ICLR 2023 · 6 citations
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- Data Acquisition via Experimental Design for Data MarketsCharles Lu, Baihe Huang, Sai Praneeth Karimireddy, Praneeth Vepakomma et al.NeurIPS 2024 · 12 citations
- DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang LowICML 2022 · 71 citations
