LAVA: Data Valuation without Pre-Specified Learning Algorithms
Hoang Anh Just, Feiyang Kang, Tianhao Wang, Yi Zeng, Myeongseob Ko, Ming Jin, Ruoxi Jia
摘要
Traditionally, data valuation is posed as a problem of equitably splitting the validation performance of a learning algorithm among the training data. As a result, the calculated data values depend on many design choices of the underlying learning algorithm. However, this dependence is undesirable for many use cases of data valuation, such as setting priorities over different data sources in a data acquisition process and informing pricing mechanisms in a data marketplace. In these scenarios, data needs to be valued before the actual analysis and the choice of the learning algorithm is still undetermined then. Another side-effect of the dependence is that to assess the value of individual points, one needs to re-run the learning algorithm with and without a point, which incurs a large computation burden. This work leapfrogs over the current limits of data valuation methods by introducing a new framework that can value training data in a way that is oblivious to the downstream learning algorithm. Our main results are as follows. (1) We develop a proxy for the validation performance associated with a training set based on a non-conventional class-wise Wasserstein distance between the training and the validation set. We show that the distance characterizes the upper bound of the validation performance for any given model under certain Lipschitz conditions. (2) We develop a novel method to value individual data based on the sensitivity analysis of the class-wise Wasserstein distance. Importantly, these values can be directly obtained for free from the output of off-the-shelf optimization solvers when computing the distance. (3) We evaluate our new data valuation framework over various use cases related to detecting low-quality data and show that, surprisingly, the learning-agnostic feature of our framework enables a significant improvement over the state-of-the-art performance while being orders of magnitude faster.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMsFeiyang Kang, Hoang Anh Just, Yifan Sun, Himanshu Jahagirdar 等ICLR 2024 · 被引用 39 次
- Stochastic Amortization: A Unified Approach to Accelerate Feature and Data AttributionIan Covert, Chanwoo Kim, Su-In Lee, James Y. Zou 等NeurIPS 2024 · 被引用 25 次
- Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon 等ICML 2024 · 被引用 24 次
- Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed SourcesFeiyang Kang, Hoang Anh Just, Anit Kumar Sahu, Ruoxi JiaNeurIPS 2023 · 被引用 21 次
- Selectivity Drives Productivity: Efficient Dataset Pruning for Enhanced Transfer LearningYihua Zhang, Yimeng Zhang, Aochuan Chen, Jinghan Jia 等NeurIPS 2023 · 被引用 18 次
它引用的顶会 Paper8
- Geometric Dataset Distances via Optimal TransportDavid Alvarez-Melis, Nicolò FusiNeurIPS 2020 · 被引用 267 次
- Unlearnable Examples: Making Personal Data UnexploitableHanxun Huang, Xingjun Ma, Sarah Monazam Erfani, James Bailey 等ICLR 2021 · 被引用 255 次
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 被引用 236 次
- Adversarial Unlearning of Backdoors via Implicit HypergradientYi Zeng, Si Chen, Won Park, Zhuoqing Mao 等ICLR 2022 · 被引用 235 次
- Narcissus: A Practical Clean-Label Backdoor Attack with Limited InformationYi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu 等CCS 2023 · 被引用 170 次
相关 Paper
- Distributionally Robust Data ValuationXiaoqiang Lin, Xinyi Xu, Zhaoxuan Wu, See-Kiong Ng 等ICML 2024 · 被引用 6 次
- KAIROS: Scalable Model-Agnostic Data ValuationJiongli Zhu, Parjanya Prashant, Alex Cloninger, Babak SalimiNeurIPS 2025 · 被引用 1 次
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 被引用 152 次
- Data Valuation and Detections in Federated LearningWenqian Li, Shuran Fu, Fengrui Zhang, Yan PangCVPR 2024
- TimeLAVA: Learning-Agnostic Valuation for Time Series DataWenqin Liu, Weizhi Quan, Aoqi Zuo, Erdun Gao 等ICML 2026
