Deep Learning Through the Lens of Example Difficulty
Robert J. N. Baldock, Hartmut Maennel, Behnam Neyshabur
摘要
Existing work on understanding deep learning often employs measures that compress all data-dependent information into a few numbers. In this work, we adopt a perspective based on the role of individual examples. We introduce a measure of the computational difficulty of making a prediction for a given input: the (effective) prediction depth. Our extensive investigation reveals surprising yet simple relationships between the prediction depth of a given input and the model's uncertainty, confidence, accuracy and speed of learning for that data point. We further categorize difficult examples into three interpretable groups, demonstrate how these groups are processed differently inside deep models and showcase how this understanding allows us to improve prediction accuracy. Insights from our study lead to a coherent view of a number of separately reported phenomena in the literature: early layers generalize while later layers memorize; early layers converge faster and networks learn easy data and simple functions first. * Work completed as part of the Google AI Residency Program 1 We expand on other notions of example difficulty in Appendix B.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper66
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 被引用 135 次
- Head2Toe: Utilizing Intermediate Representations for Better Transfer LearningUtku Evci, Vincent Dumoulin, Hugo Larochelle, Michael C. MozerICML 2022 · 被引用 103 次
- Can Neural Network Memorization Be Localized?Pratyush Maini, Michael Curtis Mozer, Hanie Sedghi, Zachary Chase Lipton 等ICML 2023 · 被引用 82 次
- Estimating Example Difficulty using Variance of GradientsChirag Agarwal, Daniel D'souza, Sara HookerCVPR 2022 · 被引用 57 次
它引用的顶会 Paper18
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 被引用 569 次
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 被引用 263 次
相关 Paper
- Uncertainty-Aware CNNs for Depth Completion: Uncertainty from Beginning to EndAbdelrahman Eldesokey, Michael Felsberg, Karl Holmquist, Michael PerssonCVPR 2020
- Your Model is Overconfident, and Other Lies We Tell OurselvesTimothee Mickus, Aman Sinha, Raúl VázquezACL 2025
- Deep Frequency Principle Towards Understanding Why Deeper Learning Is FasterZhiqin John Xu, Hanxu ZhouAAAI 2021 · 被引用 67 次
- Measuring Per-Unit Interpretability at Scale Without HumansRoland S. Zimmermann, David A. Klindt, Wieland BrendelNeurIPS 2024 · 被引用 5 次
- Understanding Visual Feature Reliance through the Lens of ComplexityThomas Fel, Louis Béthune, Andrew K. Lampinen, Thomas Serre 等NeurIPS 2024 · 被引用 20 次
