The Effect of Intrinsic Dataset Properties on Generalization: Unraveling Learning Differences Between Natural and Medical Images
Nicholas Konz, Maciej A. Mazurowski
Abstract
This paper investigates discrepancies in how neural networks learn from different imaging domains, which are commonly overlooked when adopting computer vision techniques from the domain of natural images to other specialized domains such as medical images. Recent works have found that the generalization error of a trained network typically increases with the intrinsic dimension () of its training set. Yet, the steepness of this relationship varies significantly between medical (radiological) and natural imaging domains, with no existing theoretical explanation. We address this gap in knowledge by establishing and empirically validating a generalization scaling law with respect to , and propose that the substantial scaling discrepancy between the two considered domains may be at least partially attributed to the higher intrinsic ``label sharpness'' () of medical imaging datasets, a metric which we propose. Next, we demonstrate an additional benefit of measuring the label sharpness of a training set: it is negatively correlated with the trained model's adversarial robustness, which notably leads to models for medical images having a substantially higher vulnerability to adversarial attack. Finally, we extend our formalism to the related metric of learned representation intrinsic dimension (), derive a generalization scaling law with respect to , and show that serves as an upper bound for . Our theoretical results are supported by thorough experiments with six models and eleven natural and medical imaging datasets over a range of training set sizes. Our findings offer insights into the influence of intrinsic dataset properties on generalization, representation learning, and robustness in deep neural networks. Code link: https://github.com/mazurowski-lab/intrinsic-properties
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- IdEst: Assessing Self-Supervised Learning Representations via Intrinsic DimensionJulie Mordacq, Vicky Kalogeiton, Steve OudotICML 2026 · 1 citation
- Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision TransformersUfaq Khan, Umair Nawaz, Massimo Caputo, Muhammad Bilal et al.CVPR 2026
- Adjustment for Confounding using Pre-Trained RepresentationsRickmer Schulte, David Rügamer, Thomas NaglerICML 2025
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- Intrinsic Dimension, Persistent Homology and Generalization in Neural NetworksTolga Birdal, Aaron Lou, Leonidas J. Guibas, Umut SimsekliNeurIPS 2021 · 94 citations
- Towards Certifying L-infinity Robustness using Neural Networks with L-inf-dist NeuronsBohang Zhang, Tianle Cai, Zhou Lu, Di He et al.ICML 2021 · 62 citations
Related papers
- Domain Generalization for Medical Imaging Classification with Linear-Dependency RegularizationHaoliang Li, Yufei Wang, Renjie Wan, Shiqi Wang et al.NeurIPS 2020 · 233 citations
- Explicit Tradeoffs between Adversarial and Natural Distributional RobustnessMazda Moayeri, Kiarash Banihashem, Soheil FeiziNeurIPS 2022 · 28 citations
- Towards Robust Out-of-Distribution Generalization Bounds via SharpnessYingtian Zou, Kenji Kawaguchi, Yingnan Liu, Jiashuo Liu et al.ICLR 2024 · 13 citations
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 385 citations
- A Textbook Remedy for Domain Shifts: Knowledge Priors for Medical Image AnalysisYue Yang, Mona Gandhi, Yufei Wang, Yifan Wu et al.NeurIPS 2024 · 21 citations
