Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits
Jiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon, Ruoxi Jia
Abstract
Data Shapley provides a principled approach to data valuation and plays a crucial role in data-centric machine learning (ML) research. Data selection is considered a standard application of Data Shapley. However, its data selection performance has shown to be inconsistent across settings in the literature. This study aims to deepen our understanding of this phenomenon. We introduce a hypothesis testing framework and show that Data Shapley's performance can be no better than random selection without specific constraints on utility functions. We identify a class of utility functions, monotonically transformed modular functions, within which Data Shapley optimally selects data. Based on this insight, we propose a heuristic for predicting Data Shapley's effectiveness in data selection tasks. Our experiments corroborate these findings, adding new insights into when Data Shapley may or may not succeed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53988fd8-69bf-4f1c-a7c1-303f2787f8d7Cited by top-tier papers20
- GREATS: Online Selection of High-Quality Data for LLM Training in Every IterationJiachen T. Wang, Tong Wu, Dawn Song, Prateek Mittal et al.NeurIPS 2024 · 91 citations
- Most Influential Subset Selection: Challenges, Promises, and BeyondYuzheng Hu, Pingbang Hu, Han Zhao, Jiaqi W. MaNeurIPS 2024 · 39 citations
- Influence Guided Context Selection for Effective Retrieval-Augmented GenerationJiale Deng, Yanyan Shen, Ziyuan Pei, Youmin Chen et al.NeurIPS 2025 · 8 citations
- TreeGrad-Ranker: Feature Ranking via O(L)-Time Gradients for Decision TreesWeida Li, Yaoliang Yu, Bryan Kian Hsiang LowICLR 2026 · 5 citations
- A Comprehensive Study of Shapley Value in Data AnalyticsHong Lin, Shixin Wan, Zhongle Xie, Ke Chen et al.VLDB 2025 · 4 citations
Builds on13
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 158 citations
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 152 citations
- Validation Free and Replication Robust Volume-based Data ValuationXinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang LowNeurIPS 2021 · 89 citations
Related papers
- CS-Shapley: Class-wise Shapley Values for Data Valuation in ClassificationStephanie Schoch, Haifeng Xu, Yangfeng JiNeurIPS 2022 · 56 citations
- Is Data Shapley Not Better than Random in Data Selection? Ask NASHXiao Tian, Jue Fan, Rachael Hwee Ling Sim, Zixuan Wang et al.ICML 2026
- P-Shapley: Shapley Values on Probabilistic ClassifiersHaocheng Xia, Xiang Li, Junyuan Pang, Jinfei Liu et al.VLDB 2024
- Unifying and Optimizing Data Values for Selection via Sequential Decision-MakingFrank Hongliang Chi, Qiong Wu, Zhengyi Zhou, Jonathan Li et al.ICML 2026 · 1 citation
- 2D-Shapley: A Framework for Fragmented Data ValuationZhihong Liu, Hoang Anh Just, Xiangyu Chang, Xi Chen et al.ICML 2023 · 13 citations
