Verifying the Union of Manifolds Hypothesis for Image Data
Bradley C. A. Brown, Anthony L. Caterini, Brendan Leigh Ross, Jesse C. Cresswell, Gabriel Loaiza-Ganem
Abstract
Deep learning has had tremendous success at learning low-dimensional representations of high-dimensional data. This success would be impossible if there was no hidden low-dimensional structure in data of interest; this existence is posited by the manifold hypothesis, which states that the data lies on an unknown manifold of low intrinsic dimension. In this paper, we argue that this hypothesis does not properly capture the low-dimensional structure typically present in image data. Assuming that data lies on a single manifold implies intrinsic dimension is identical across the entire data space, and does not allow for subregions of this space to have a different number of factors of variation. To address this deficiency, we consider the union of manifolds hypothesis, which states that data lies on a disjoint union of manifolds of varying intrinsic dimensions. We empirically verify this hypothesis on commonly-used image datasets, finding that indeed, observed data lies on a disconnected set and that intrinsic dimension is not constant. We also provide insights into the implications of the union of manifolds hypothesis in deep learning, both supervised and unsupervised, showing that designing models with an inductive bias for this structure improves performance across classification and generative modelling tasks. Our code is available at https://github.com/layer6ai-labs/UoMH.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7bac167-2ec8-412a-a14d-2e69463851e4Cited by top-tier papers39
- Guiding a Diffusion Model with a Bad Version of ItselfTero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen et al.NeurIPS 2024 · 338 citations
- A Geometric View of Data Complexity: Efficient Local Intrinsic Dimension Estimation with Diffusion ModelsHamidreza Kamkari, Brendan Leigh Ross, Rasa Hosseinzadeh, Jesse C. Cresswell et al.NeurIPS 2024 · 49 citations
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion ModelsChubin Chen, Jiashu Zhu, Xiaokun Feng, Nisha Huang et al.ICLR 2026 · 44 citations
- Hardness of Learning Neural Networks under the Manifold HypothesisBobak T. Kiani, Jason Wang, Melanie WeberNeurIPS 2024 · 25 citations
- A Geometric Explanation of the Likelihood OOD Detection ParadoxHamidreza Kamkari, Brendan Leigh Ross, Jesse C. Cresswell, Anthony L. Caterini et al.ICML 2024 · 20 citations
Builds on15
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- Generalized Energy Based ModelsMichael Arbel, Liang Zhou, Arthur GrettonICLR 2021 · 254 citations
- Riemannian Continuous Normalizing FlowsEmile Mathieu, Maximilian NickelNeurIPS 2020 · 198 citations
- Flows for simultaneous manifold learning and density estimationJohann Brehmer, Kyle CranmerNeurIPS 2020 · 187 citations
- Normalizing Flows on Tori and SpheresDanilo Jimenez Rezende, George Papamakarios, Sébastien Racanière, Michael S. Albergo et al.ICML 2020 · 181 citations
Related papers
- CW Complex Hypothesis for Image DataYi Wang, Zhiren WangICML 2024 · 2 citations
- On Deep Generative Models for Approximation and Estimation of Distributions on ManifoldsBiraj Dahal, Alexander Havrilla, Minshuo Chen, Tuo Zhao et al.NeurIPS 2022 · 17 citations
- Data Representations' Study of Latent Image ManifoldsIlya Kaufman, Omri AzencotICML 2023 · 11 citations
- A Critique of Self-Expressive Deep Subspace ClusteringBenjamin David Haeffele, Chong You, René VidalICLR 2021 · 35 citations
- Topological Singularity Detection at Multiple ScalesJulius von Rohrscheidt, Bastian RieckICML 2023 · 16 citations
