Two Calm Ends and the Wild Middle: A Geometric Picture of Memorization in Diffusion Models
Nick Dodson, Xinyu Gao, Qingsong Wang, Yusu Wang, Zhengchao Wan
摘要
Diffusion models generate high-quality samples but can also memorize training data, raising serious privacy concerns. Understanding the mechanisms governing when memorization versus generalization occurs remains an active area of research. In particular, it is unclear where along the noise schedule memorization is induced, how data geometry influences it, and how phenomena at different noise scales interact. We introduce a geometric framework that partitions the noise schedule into three regimes based on the coverage properties of training data by Gaussian shells and the concentration behavior of the posterior, which we argue are two fundamental objects governing memorization and generalization in diffusion models. This perspective reveals that memorization risk is highly non-uniform across noise levels. We further identify a danger zone at medium noise levels where memorization is most pronounced. In contrast, both the small and large noise regimes resist memorization, but through fundamentally different mechanisms: small noise avoids memorization due to limited training coverage, while large noise exhibits low posterior concentration and admits a provably near linear Gaussian denoising behavior. For the medium noise regime, we identify geometric conditions through which we propose a geometry-informed targeted intervention that mitigates memorization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Generalization in diffusion models arises from geometry-adaptive harmonic representationsZahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, Stéphane MallatICLR 2024 · 被引用 168 次
- Detecting, Explaining, and Mitigating Memorization in Diffusion ModelsYuxin Wen, Yuchen Liu, Chen Chen, Lingjuan LyuICLR 2024 · 被引用 103 次
- Why Diffusion Models Don't Memorize: The Role of Implicit Dynamical Regularization in TrainingTony Bonnaire, Raphaël Urfin, Giulio Biroli, Marc MézardNeurIPS 2025 · 被引用 93 次
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel 等ICLR 2023 · 被引用 87 次
相关 Paper
- Does Generation Require Memorization? Creative Diffusion Models using Ambient DiffusionKulin Shah, Alkis Kalavasis, Adam R. Klivans, Giannis DarasICML 2025
- On the Edge of Memorization in Diffusion ModelsSam Buchanan, Druv Pai, Yi Ma, Valentin De BortoliNeurIPS 2025 · 被引用 25 次
- The Emergence of Reproducibility and Consistency in Diffusion ModelsHuijie Zhang, Jinfan Zhou, Yifu Lu, Minzhe Guo 等ICML 2024 · 被引用 51 次
- Understanding Generalizability of Diffusion Models Requires Rethinking the Hidden Gaussian StructureXiang Li, Yixiang Dai, Qing QuNeurIPS 2024 · 被引用 45 次
- Latent Diffusion Inversion Requires Understanding the Latent SpaceMingxing Rao, Bowen Qu, Daniel MoyerCVPR 2026 · 被引用 4 次
