Data Valuation Without Training of a Model
Nohyun Ki, Hoyong Choi, Hye Won Chung
摘要
Many recent works on understanding deep learning try to quantify how much individual data instances influence the optimization and generalization of a model. Such attempts reveal characteristics and importance of individual instances, which may provide useful information in diagnosing and improving deep learning. However, most of the existing works on data valuation require actual training of a model, which often demands high-computational cost. In this paper, we provide a training-free data valuation score, called complexity-gap score, which is a datacentric score to quantify the influence of individual instances in generalization of two-layer overparameterized neural networks. The proposed score can quantify irregularity of the instances and measure how much each data instance contributes in the total movement of the network parameters during training. We theoretically analyze and empirically demonstrate the effectiveness of the complexity-gap score in finding 'irregular or mislabeled' data instances, and also provide applications of the score in analyzing datasets and diagnosing training dynamics. Our code is publicly available at https://github.com/JJchy/CG_score .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- DataInf: Efficiently Estimating Data Influence in LoRA-tuned LLMs and Diffusion ModelsYongchan Kwon, Eric Wu, Kevin Wu, James ZouICLR 2024 · 被引用 112 次
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao 等NeurIPS 2025 · 被引用 112 次
- Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon 等ICML 2024 · 被引用 24 次
- BWS: Best Window Selection Based on Sample Scores for Data Pruning across Broad RangesHoyong Choi, Nohyun Ki, Hye Won ChungICML 2024 · 被引用 9 次
- Distributionally Robust Data ValuationXiaoqiang Lin, Xinyi Xu, Zhaoxuan Wu, See-Kiong Ng 等ICML 2024 · 被引用 6 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
相关 Paper
- DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang LowICML 2022 · 被引用 71 次
- Detecting Corrupted Labels Without Training a Model to PredictZhaowei Zhu, Zihao Dong, Yang LiuICML 2022 · 被引用 84 次
- HyDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural NetworksYuanyuan Chen, Boyang Li, Han Yu, Pengcheng Wu 等AAAI 2021 · 被引用 50 次
- Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification TasksFlorian A. Hölzl, Daniel Rueckert, Georgios KaissisNeurIPS 2025 · 被引用 1 次
- Inconsistency-Aware Minimization: Improving Generalization with Unlabeled DataHee-Sung Kim, Hyeonseong Kim, Sungyoon LeeICML 2026
