A Geometric Framework for Understanding Memorization in Generative Models
Brendan Leigh Ross, Hamidreza Kamkari, Tongzi Wu, Rasa Hosseinzadeh, Zhaoyan Liu, George Stein, Jesse C. Cresswell, Gabriel Loaiza-Ganem
Abstract
As deep generative models have progressed, recent work has shown that they are capable of memorizing and reproducing training datapoints when deployed. These findings call into question the usability of generative models, especially in light of the legal and privacy risks brought about by memorization. To better understand this phenomenon, we propose a geometric framework which leverages the manifold hypothesis into a clear language in which to reason about memorization. We propose to analyze memorization in terms of the relationship between the dimensionalities of (i) the ground truth data manifold and (ii) the manifold learned by the model. In preliminary tests on toy examples and Stable Diffusion (Rombach et al., 2022) , we show that our theoretical framework accurately describes reality. Furthermore, by analyzing prior work in the context of our geometric framework, we explain and unify assorted observations in the literature and illuminate promising directions for future research on memorization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 044d2279-947e-4ad2-8907-e0765af61e95Cited by top-tier papers17
- On the Closed-Form of Flow Matching: Generalization Does Not Arise from Target StochasticityQuentin Bertrand, Anne Gagneux, Mathurin Massias, Rémi EmonetNeurIPS 2025 · 49 citations
- On the Edge of Memorization in Diffusion ModelsSam Buchanan, Druv Pai, Yi Ma, Valentin De BortoliNeurIPS 2025 · 25 citations
- A Closer Look at Model Collapse: From a Generalization-to-Memorization PerspectiveLianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang et al.NeurIPS 2025 · 22 citations
- Provable Separations between Memorization and Generalization in Diffusion ModelsZeqi Ye, Qijie Zhu, Molei Tao, Minshuo ChenICLR 2026 · 15 citations
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi et al.ICLR 2026 · 14 citations
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Score-based Generative Modeling in Latent SpaceArash Vahdat, Karsten Kreis, Jan KautzNeurIPS 2021 · 903 citations
Related papers
- On the geometry of generalization and memorization in deep neural networksCory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui et al.ICLR 2021 · 95 citations
- Localizing Memorized Regions in Diffusion Models via Coordinate-Wise Curvature DifferencesGwangho Kim, Sungyoon LeeICML 2026
- The Data Manifold under the MicroscopeMarios Koulakis, Constantin SeiboldICML 2026
- Latent Diffusion Inversion Requires Understanding the Latent SpaceMingxing Rao, Bowen Qu, Daniel MoyerCVPR 2026 · 4 citations
- Two Calm Ends and the Wild Middle: A Geometric Picture of Memorization in Diffusion ModelsNick Dodson, Xinyu Gao, Qingsong Wang, Yusu Wang et al.ICML 2026 · 4 citations
