A Geometric View of Data Complexity: Efficient Local Intrinsic Dimension Estimation with Diffusion Models
Hamidreza Kamkari, Brendan Leigh Ross, Rasa Hosseinzadeh, Jesse C. Cresswell, Gabriel Loaiza-Ganem
摘要
High-dimensional data commonly lies on low-dimensional submanifolds, and estimating the local intrinsic dimension (LID) of a datum -- i.e. the dimension of the submanifold it belongs to -- is a longstanding problem. LID can be understood as the number of local factors of variation: the more factors of variation a datum has, the more complex it tends to be. Estimating this quantity has proven useful in contexts ranging from generalization in neural networks to detection of out-of-distribution data, adversarial examples, and AI-generated text. The recent successes of deep generative models present an opportunity to leverage them for LID estimation, but current methods based on generative models produce inaccurate estimates, require more than a single pre-trained model, are computationally intensive, or do not exploit the best available deep generative models: diffusion models (DMs). In this work, we show that the Fokker-Planck equation associated with a DM can provide an LID estimator which addresses the aforementioned deficiencies. Our estimator, called FLIPD, is easy to implement and compatible with all popular DMs. Applying FLIPD to synthetic LID estimation benchmarks, we find that DMs implemented as fully-connected networks are highly effective LID estimators that outperform existing baselines. We also apply FLIPD to natural images where the true LID is unknown. Despite being sensitive to the choice of network architecture, FLIPD estimates remain a useful measure of relative complexity; compared to competing estimators, FLIPD exhibits a consistently higher correlation with image PNG compression rate and better aligns with qualitative assessments of complexity. Notably, FLIPD is orders of magnitude faster than other LID estimators, and the first to be tractable at the scale of Stable Diffusion.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Energy Matching: Unifying Flow Matching and Energy-Based Models for Generative ModelingMichal Balcerak, Tamaz Amiranashvili, Antonio Terpin, Suprosanna Shit 等NeurIPS 2025 · 被引用 33 次
- Learning normalized image densities via dual score matchingFlorentin Guth, Zahra Kadkhodaie, Eero P. SimoncelliNeurIPS 2025 · 被引用 22 次
- Diffusion Models and the Manifold Hypothesis: Log-Domain Smoothing is Geometry AdaptiveTyler Farghly, Peter Potaptchik, Samuel Howard, George Deligiannidis 等NeurIPS 2025 · 被引用 18 次
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi 等ICLR 2026 · 被引用 14 次
- Shallow Diffuse: Robust and Invisible Watermarking through Low-Dim Subspaces in Diffusion ModelsWenda Li, Huijie Zhang, Qing QuNeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 被引用 958 次
- Score-based Generative Modeling in Latent SpaceArash Vahdat, Karsten Kreis, Jan KautzNeurIPS 2021 · 被引用 903 次
相关 Paper
- A Geometric Explanation of the Likelihood OOD Detection ParadoxHamidreza Kamkari, Brendan Leigh Ross, Jesse C. Cresswell, Anthony L. Caterini 等ICML 2024 · 被引用 20 次
- A Wiener Process Perspective on Local Intrinsic Dimension Estimation MethodsPiotr Tempczyk, Lukasz Garncarek, Dominik Filipiak, Adam KurpiszAAAI 2025 · 被引用 2 次
- LIDL: Local Intrinsic Dimension Estimation Using Approximate LikelihoodPiotr Tempczyk, Rafal Michaluk, Lukasz Garncarek, Przemyslaw Spurek 等ICML 2022 · 被引用 39 次
- Why We Need New Benchmarks for Local Intrinsic Dimension EstimationPiotr Tempczyk, Dominik Filipiak, Lukasz Garncarek, Ksawery Smoczynski 等ICLR 2026
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
