The Underlying Universal Statistical Structure of Natural Datasets
Noam Itzhak Levi, Yaron Oz
摘要
We study universal properties in real-world complex and synthetically generated datasets. Our approach is to analogize data to a physical system and employ tools from statistical physics and Random Matrix Theory (RMT) to reveal their underlying structure. Examining the local and global eigenvalue statistics of feature-feature covariance matrices, we find: (i) bulk eigenvalue power-law scaling vastly differs between uncorrelated Gaussian and real-world data, (ii) this power law behavior is reproducible using Gaussian data with long-range correlations, (iii) all dataset types exhibit chaotic RMT universality, (iv) RMT statistics emerge at smaller dataset sizes than typical training sets, correlating with power-law convergence, (v) Shannon entropy correlates with RMT structure and requires fewer samples in strongly correlated datasets. These results suggest natural image Gram matrices can be approximated by Wishart random matrices with simple covariance structure, enabling rigorous analysis of neural network behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard 等ICML 2020 · 被引用 184 次
- Revisiting Neural Scaling Laws in Language and VisionIbrahim M. Alabdulmohsin, Behnam Neyshabur, Xiaohua ZhaiNeurIPS 2022 · 被引用 171 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
相关 Paper
- Analyzing Neural Scaling Laws in Two-Layer Networks with Power-Law Data SpectraRoman Worschech, Bernd RosenowICLR 2025
- Cascade of phase transitions in the training of energy-based modelsDimitrios Bachtis, Giulio Biroli, Aurélien Decelle, Beatriz SeoaneNeurIPS 2024 · 被引用 17 次
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 被引用 90 次
- Models of Heavy-Tailed Mechanistic UniversalityLiam Hodgkinson, Zhichao Wang, Michael W. MahoneyICML 2025
- Random Matrix Theory Proves that Deep Learning Representations of GAN-data Behave as Gaussian MixturesMohamed El Amine Seddik, Cosme Louart, Mohamed Tamaazousti, Romain CouilletICML 2020 · 被引用 78 次
