The Data Manifold under the Microscope
Marios Koulakis, Constantin Seibold
摘要
A significant gap exists between theory and practice in deep learning. Generalization and approximation error bounds are often derived for simplified models or are too loose to be informative. Many rely on the manifold hypothesis and on geometric regularity such as intrinsic dimension, curvature, and reach. Progress requires insight into data-manifold geometry and suitable benchmarks, yet existing options are polarized: analytic manifolds with known geometry but limited applicability, or real-world datasets where geometry is only coarsely estimable. We introduce a benchmarking framework for studying data geometry. We repurpose and extend dSprites and COIL-20 with additional transformation dimensions and dense, axis-aligned sampling, and pair them with finite-difference estimators that recover curvature, reach, and volume at near-ground-truth accuracy in a regime where general-purpose estimators are unreliable or difficult to deploy. The framework is intended as a controlled testbed, useful as a calibration environment for geometric estimators and a sandbox for probing theoretical assumptions. To illustrate its use, we present two application studies, namely assessing the scaling behavior of the bounds of Genovese et al. and Fefferman et al., and tracking the layer-wise geometry of a -VAE, highlighting the behavior of current bounds and the value of controlled benchmarks for guiding and validating future theory. A reference implementation is available at https://github.com/koulakis/manifold-microscope.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
- Intrinsic Dimension, Persistent Homology and Generalization in Neural NetworksTolga Birdal, Aaron Lou, Leonidas J. Guibas, Umut SimsekliNeurIPS 2021 · 被引用 94 次
- Denoising Normalizing FlowChristian Horvat, Jean-Pascal PfisterNeurIPS 2021 · 被引用 39 次
相关 Paper
- A Geometric Framework for Understanding Memorization in Generative ModelsBrendan Leigh Ross, Hamidreza Kamkari, Tongzi Wu, Rasa Hosseinzadeh 等ICLR 2025
- Why We Need New Benchmarks for Local Intrinsic Dimension EstimationPiotr Tempczyk, Dominik Filipiak, Lukasz Garncarek, Ksawery Smoczynski 等ICLR 2026
- Data Representations' Study of Latent Image ManifoldsIlya Kaufman, Omri AzencotICML 2023 · 被引用 11 次
- Carré du champ flow matching: better quality-generalisation tradeoff in generative modelsJacob Bamberger, Iolo Jones, Dennis Duncan, Michael M. Bronstein 等ICLR 2026 · 被引用 14 次
- On Deep Generative Models for Approximation and Estimation of Distributions on ManifoldsBiraj Dahal, Alexander Havrilla, Minshuo Chen, Tuo Zhao 等NeurIPS 2022 · 被引用 17 次
