Estimating Example Difficulty using Variance of Gradients
Chirag Agarwal, Daniel D'souza, Sara Hooker
摘要
In machine learning, a question of great interest is understanding what examples are challenging for a model to classify. Identifying atypical examples ensures the safe de-ployment of models, isolates samples that require further human inspection and provides interpretability into model behavior. In this work, we propose Variance of Gradients (VoG) as a valuable and efficient metric to rank data by difficulty and to surface a tractable subset of the most chal-lenging examples for human-in-the-loop auditing. We show that data points with high VoG scores are far more difficult for the model to learn and over-index on corrupted or mem-orized examples. Further, restricting the evaluation to the test set instances with the lowest VoG improves the model's generalization performance. Finally, we show that VoG is a valuable and efficient ranking for out-of-distribution detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 被引用 204 次
- Combining Ensembles and Data Augmentation Can Harm Your CalibrationYeming Wen, Ghassen Jerfel, Rafael Muller, Michael W. Dusenberry 等ICLR 2021 · 被引用 72 次
- Trivial or Impossible --- dichotomous data difficulty masks model differences (on ImageNet and beyond)Kristof Meding, Luca M. Schulze Buschoff, Robert Geirhos, Felix A. WichmannICLR 2022 · 被引用 47 次
- Generating High Fidelity Data from Low-density Regions using Diffusion ModelsVikash Sehwag, Caner Hazirbas, Albert Gordo, Firat Ozgenel 等CVPR 2022 · 被引用 38 次
它引用的顶会 Paper7
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Deep Semi-Supervised Anomaly DetectionLukas Ruff, Robert A. Vandermeulen, Nico Görnitz, Alexander Binder 等ICLR 2020 · 被引用 678 次
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 被引用 204 次
- Characterizing Structural Regularities of Labeled Data in Overparameterized ModelsZiheng Jiang, Chiyuan Zhang, Kunal Talwar, Michael C. MozerICML 2021 · 被引用 128 次
- Does learning require memorization? a short tale about a long tailVitaly FeldmanSTOC 2020 · 被引用 28 次
相关 Paper
- GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution DetectionJinggang Chen, Junjie Li, Xiaoyang Qu, Jianzong Wang 等NeurIPS 2023 · 被引用 17 次
- Influence Scores at Scale for Efficient Language Data SamplingNikhil Anand, Joshua Tan, Maria MinakovaEMNLP 2023 · 被引用 1 次
- Out-Of-Distribution Detection with Diversification (Provably)Haiyun Yao, Zongbo Han, Huazhu Fu, Xi Peng 等NeurIPS 2024 · 被引用 9 次
- GradOrth: A Simple yet Efficient Out-of-Distribution Detection with Orthogonal Projection of GradientsSima Behpour, Thang Long Doan, Xin Li, Wenbin He 等NeurIPS 2023 · 被引用 34 次
- Leveraging Perturbation Robustness to Enhance Out-of-Distribution DetectionWenxi Chen, Raymond A. Yeh, Shaoshuai Mou, Yan GuCVPR 2025
