EcoVal: An Efficient Data Valuation Framework for Machine Learning
Ayush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan, Bowei Chen, Mohan S. Kankanhalli
摘要
Quantifying the value of data within a machine learning workflow can play a pivotal role in making more strategic decisions in machine learning initiatives. The existing Shapley value based frameworks for data valuation in machine learning are computationally expensive as they require considerable amount of repeated training of the model to obtain the Shapley value. In this paper, we introduce an efficient data valuation framework EcoVal, to estimate the value of data for machine learning models in a fast and practical manner. Instead of directly working with individual data sample, we determine the value of a cluster of similar data points. This value is further propagated amongst all the member cluster points. We show that the overall value of the data can be determined by estimating the intrinsic and extrinsic value of each data. This is enabled by formulating the performance of a model as aproduction function, a concept which is popularly used to estimate the amount of output based on factors like labor and capital in a traditional free economic market. We provide a formal proof of our valuation technique and elucidate the principles and mechanisms that enable its accelerated performance. We demonstrate the real-world applicability of our method by showcasing its effectiveness for both in-distribution and out-of-sample data. This work addresses one of the core challenges of efficient data valuation at scale in machine learning models. The code is available at https://github.com/respai-lab/ecoval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- The Many Shapley Values for Model ExplanationMukund Sundararajan, Amir NajmiICML 2020 · 被引用 799 次
- Neuron Shapley: Discovering the Responsible NeuronsAmirata Ghorbani, James Y. ZouNeurIPS 2020 · 被引用 160 次
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 被引用 152 次
- Datamodels: Understanding Predictions with Data and Data with PredictionsAndrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc 等ICML 2022 · 被引用 66 次
- Data Valuation Without Training of a ModelNohyun Ki, Hoyong Choi, Hye Won ChungICLR 2023 · 被引用 7 次
相关 Paper
- Localized Data Shapley: Accelerating Valuation for Nearest Neighbor AlgorithmsGuangyi Zhang, Yanhao Wang, Chengliang Chai, Qiyu Liu 等NeurIPS 2025 · 被引用 1 次
- CS-Shapley: Class-wise Shapley Values for Data Valuation in ClassificationStephanie Schoch, Haifeng Xu, Yangfeng JiNeurIPS 2022 · 被引用 56 次
- Shapley-Based Data Valuation for Weighted -Nearest NeighborsGuangyi Zhang, Qiyu Liu, Aristides GionisNeurIPS 2025 · 被引用 2 次
- Data-OOB: Out-of-bag Estimate as a Simple and Efficient Data ValueYongchan Kwon, James ZouICML 2023 · 被引用 54 次
- Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling AlgorithmBoxin Zhao, Boxiang Lyu, Raul Castro Fernandez, Mladen KolarICML 2023 · 被引用 14 次
