The Effect of Intrinsic Dataset Properties on Generalization: Unraveling Learning Differences Between Natural and Medical Images
Nicholas Konz, Maciej A. Mazurowski
摘要
This paper investigates discrepancies in how neural networks learn from different imaging domains, which are commonly overlooked when adopting computer vision techniques from the domain of natural images to other specialized domains such as medical images. Recent works have found that the generalization error of a trained network typically increases with the intrinsic dimension () of its training set. Yet, the steepness of this relationship varies significantly between medical (radiological) and natural imaging domains, with no existing theoretical explanation. We address this gap in knowledge by establishing and empirically validating a generalization scaling law with respect to , and propose that the substantial scaling discrepancy between the two considered domains may be at least partially attributed to the higher intrinsic ``label sharpness'' () of medical imaging datasets, a metric which we propose. Next, we demonstrate an additional benefit of measuring the label sharpness of a training set: it is negatively correlated with the trained model's adversarial robustness, which notably leads to models for medical images having a substantially higher vulnerability to adversarial attack. Finally, we extend our formalism to the related metric of learned representation intrinsic dimension (), derive a generalization scaling law with respect to , and show that serves as an upper bound for . Our theoretical results are supported by thorough experiments with six models and eleven natural and medical imaging datasets over a range of training set sizes. Our findings offer insights into the influence of intrinsic dataset properties on generalization, representation learning, and robustness in deep neural networks. Code link: https://github.com/mazurowski-lab/intrinsic-properties
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- IdEst: Assessing Self-Supervised Learning Representations via Intrinsic DimensionJulie Mordacq, Vicky Kalogeiton, Steve OudotICML 2026 · 被引用 1 次
- Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision TransformersUfaq Khan, Umair Nawaz, Massimo Caputo, Muhammad Bilal 等CVPR 2026
- Adjustment for Confounding using Pre-Trained RepresentationsRickmer Schulte, David Rügamer, Thomas NaglerICML 2025
它引用的顶会 Paper8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
- Intrinsic Dimension, Persistent Homology and Generalization in Neural NetworksTolga Birdal, Aaron Lou, Leonidas J. Guibas, Umut SimsekliNeurIPS 2021 · 被引用 94 次
- Towards Certifying L-infinity Robustness using Neural Networks with L-inf-dist NeuronsBohang Zhang, Tianle Cai, Zhou Lu, Di He 等ICML 2021 · 被引用 62 次
相关 Paper
- Domain Generalization for Medical Imaging Classification with Linear-Dependency RegularizationHaoliang Li, Yufei Wang, Renjie Wan, Shiqi Wang 等NeurIPS 2020 · 被引用 233 次
- Explicit Tradeoffs between Adversarial and Natural Distributional RobustnessMazda Moayeri, Kiarash Banihashem, Soheil FeiziNeurIPS 2022 · 被引用 28 次
- Towards Robust Out-of-Distribution Generalization Bounds via SharpnessYingtian Zou, Kenji Kawaguchi, Yingnan Liu, Jiashuo Liu 等ICLR 2024 · 被引用 13 次
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 被引用 385 次
- A Textbook Remedy for Domain Shifts: Knowledge Priors for Medical Image AnalysisYue Yang, Mona Gandhi, Yufei Wang, Yifan Wu 等NeurIPS 2024 · 被引用 21 次
