Data Valuation using Reinforcement Learning
Jinsung Yoon, Sercan Ömer Arik, Tomas Pfister
摘要
Quantifying the value of data is a fundamental problem in machine learning. Data valuation has multiple important use cases: (1) building insights about the learning task, (2) domain adaptation, (3) corrupted sample discovery, and (4) robust learning. To adaptively learn data values jointly with the target task predictor model, we propose a meta learning framework which we name Data Valuation using Reinforcement Learning (DVRL). We employ a data value estimator (modeled by a deep neural network) to learn how likely each datum is used in training of the predictor model. We train the data value estimator using a reinforcement signal of the reward obtained on a small validation set that reflects performance on the target task. We demonstrate that DVRL yields superior data value estimates compared to alternative methods across different types of datasets and in a diverse set of application scenarios. The corrupted sample discovery performance of DVRL is close to optimal in many regimes (i.e. as if the noisy samples were known apriori), and for domain adaptation and robust learning DVRL significantly outperforms state-of-the-art by 14.6% and 10.8%, respectively. 1. Incorrect label (e.g. human labeling errors). 2. Input comes from a different distribution (e.g. different location or time). 3. Input is noisy or low quality (e.g. noisy capturing hardware). 4. Usefulness for target task (label is very common in the training dataset but not as common in the testing dataset). In addition to improving performance in such scenarios, data valuation also enables many new use cases. It can suggest better practices for data collection, e.g. what kinds of additional data would the *Work done as an intern at Google Cloud.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 被引用 158 次
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 被引用 112 次
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao 等NeurIPS 2025 · 被引用 112 次
- Learning Fast Sample Re-weighting Without Reward DataZizhao Zhang, Tomas PfisterICCV 2021 · 被引用 109 次
- Sensitivity-Aware Visual Parameter-Efficient Fine-TuningHaoyu He, Jianfei Cai, Jing Zhang, Dacheng Tao 等ICCV 2023 · 被引用 97 次
相关 Paper
- Distributionally Robust Data ValuationXiaoqiang Lin, Xinyi Xu, Zhaoxuan Wu, See-Kiong Ng 等ICML 2024 · 被引用 6 次
- Multi-Source Deep Domain Adaptation with Weak Supervision for Time-Series Sensor DataGarrett Wilson, Janardhan Rao Doppa, Diane J. CookKDD 2020 · 被引用 6 次
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution CorrectionAviral Kumar, Abhishek Gupta, Sergey LevineNeurIPS 2020 · 被引用 124 次
- Doubly Robust Augmented Transfer for Meta-Reinforcement LearningYuankun Jiang, Nuowen Kan, Chenglin Li, Wenrui Dai 等NeurIPS 2023 · 被引用 3 次
- DeRDaVa: Deletion-Robust Data Valuation for Machine LearningXiao Tian, Rachael Hwee Ling Sim, Jue Fan, Bryan Kian Hsiang LowAAAI 2024 · 被引用 3 次
