Verifying the Union of Manifolds Hypothesis for Image Data
Bradley C. A. Brown, Anthony L. Caterini, Brendan Leigh Ross, Jesse C. Cresswell, Gabriel Loaiza-Ganem
摘要
Deep learning has had tremendous success at learning low-dimensional representations of high-dimensional data. This success would be impossible if there was no hidden low-dimensional structure in data of interest; this existence is posited by the manifold hypothesis, which states that the data lies on an unknown manifold of low intrinsic dimension. In this paper, we argue that this hypothesis does not properly capture the low-dimensional structure typically present in image data. Assuming that data lies on a single manifold implies intrinsic dimension is identical across the entire data space, and does not allow for subregions of this space to have a different number of factors of variation. To address this deficiency, we consider the union of manifolds hypothesis, which states that data lies on a disjoint union of manifolds of varying intrinsic dimensions. We empirically verify this hypothesis on commonly-used image datasets, finding that indeed, observed data lies on a disconnected set and that intrinsic dimension is not constant. We also provide insights into the implications of the union of manifolds hypothesis in deep learning, both supervised and unsupervised, showing that designing models with an inductive bias for this structure improves performance across classification and generative modelling tasks. Our code is available at https://github.com/layer6ai-labs/UoMH.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Guiding a Diffusion Model with a Bad Version of ItselfTero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen 等NeurIPS 2024 · 被引用 338 次
- A Geometric View of Data Complexity: Efficient Local Intrinsic Dimension Estimation with Diffusion ModelsHamidreza Kamkari, Brendan Leigh Ross, Rasa Hosseinzadeh, Jesse C. Cresswell 等NeurIPS 2024 · 被引用 49 次
- Stochastic Self-Guidance for Training-Free Enhancement of Diffusion ModelsChubin Chen, Jiashu Zhu, Xiaokun Feng, Nisha Huang 等ICLR 2026 · 被引用 44 次
- Hardness of Learning Neural Networks under the Manifold HypothesisBobak T. Kiani, Jason Wang, Melanie WeberNeurIPS 2024 · 被引用 25 次
- A Geometric Explanation of the Likelihood OOD Detection ParadoxHamidreza Kamkari, Brendan Leigh Ross, Jesse C. Cresswell, Anthony L. Caterini 等ICML 2024 · 被引用 20 次
它引用的顶会 Paper15
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
- Generalized Energy Based ModelsMichael Arbel, Liang Zhou, Arthur GrettonICLR 2021 · 被引用 254 次
- Riemannian Continuous Normalizing FlowsEmile Mathieu, Maximilian NickelNeurIPS 2020 · 被引用 198 次
- Flows for simultaneous manifold learning and density estimationJohann Brehmer, Kyle CranmerNeurIPS 2020 · 被引用 187 次
- Normalizing Flows on Tori and SpheresDanilo Jimenez Rezende, George Papamakarios, Sébastien Racanière, Michael S. Albergo 等ICML 2020 · 被引用 181 次
相关 Paper
- CW Complex Hypothesis for Image DataYi Wang, Zhiren WangICML 2024 · 被引用 2 次
- On Deep Generative Models for Approximation and Estimation of Distributions on ManifoldsBiraj Dahal, Alexander Havrilla, Minshuo Chen, Tuo Zhao 等NeurIPS 2022 · 被引用 17 次
- Data Representations' Study of Latent Image ManifoldsIlya Kaufman, Omri AzencotICML 2023 · 被引用 11 次
- A Critique of Self-Expressive Deep Subspace ClusteringBenjamin David Haeffele, Chong You, René VidalICLR 2021 · 被引用 35 次
- Topological Singularity Detection at Multiple ScalesJulius von Rohrscheidt, Bastian RieckICML 2023 · 被引用 16 次
