On Memorization in Probabilistic Deep Generative Models
Gerrit J. J. van den Burg, Christopher K. I. Williams
Abstract
Recent advances in deep generative models have led to impressive results in a variety of application domains. Motivated by the possibility that deep learning models might memorize part of the input data, there have been increased efforts to understand how memorization arises. In this work, we extend a recently proposed measure of memorization for supervised learning (Feldman, 2019) to the unsupervised density estimation problem and adapt it to be more computationally efficient. Next, we present a study that demonstrates how memorization can occur in probabilistic deep generative models such as variational autoencoders. This reveals that the form of memorization to which these models are susceptible differs fundamentally from mode collapse and overfitting. Furthermore, we show that the proposed memorization score measures a phenomenon that is not captured by commonly-used nearest neighbor tests. Finally, we discuss several strategies that can be used to limit memorization in practice. Our work thus provides a framework for understanding problematic memorization in probabilistic generative models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85df45d0-59a0-43d6-9c26-72bbce7d54d3Cited by top-tier papers16
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- Enhanced Membership Inference Attacks against Machine Learning ModelsJiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler et al.CCS 2022 · 150 citations
- Structure-informed Language Models Are Protein DesignersZaixiang Zheng, Yifan Deng, Dongyu Xue, Yi Zhou et al.ICML 2023 · 130 citations
- Feature Likelihood Score: Evaluating the Generalization of Generative Models Using SamplesMarco Jiralerspong, Avishek Joey Bose, Ian Gemp, Chongli Qin et al.NeurIPS 2023 · 39 citations
- Understanding Instance-based Interpretability of Variational Auto-EncodersZhifeng Kong, Kamalika ChaudhuriNeurIPS 2021 · 32 citations
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Deep Models Under the GAN: Information Leakage from Collaborative Deep LearningBriland Hitaj, Giuseppe Ateniese, Fernando Pérez-CruzCCS 2017 · 1,581 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Unveiling Privacy, Memorization, and Input Curvature LinksDeepak Ravikumar, Efstathia Soufleri, Abolfazl Hashemi, Kaushik RoyICML 2024 · 16 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Memorization Through the Lens of Curvature of Loss Function Around SamplesIsha Garg, Deepak Ravikumar, Kaushik RoyICML 2024 · 25 citations
- Posterior Collapse of a Linear Latent Variable ModelZihao Wang, Liu ZiyinNeurIPS 2022 · 29 citations
- A Bayesian Nonparametrics View into Deep RepresentationsMichal Jamroz, Marcin Kurdziel, Mateusz OpalaNeurIPS 2020 · 1 citation
