Memorization in Self-Supervised Learning Improves Downstream Generalization
Wenhao Wang, Muhammad Ahmad Kaleem, Adam Dziedzic, Michael Backes, Nicolas Papernot, Franziska Boenisch
摘要
Self-supervised learning (SSL) has recently received significant attention due to its ability to train high-performance encoders purely on unlabeled data-often scraped from the internet. This data can still be sensitive and empirical evidence suggests that SSL encoders memorize private information of their training data and can disclose them at inference time. Since existing theoretical definitions of memorization from supervised learning rely on labels, they do not transfer to SSL. To address this gap, we propose SSLMem, a framework for defining memorization within SSL. Our definition compares the difference in alignment of representations for data points and their augmented views returned by both encoders that were trained on these data points and encoders that were not. Through comprehensive empirical analysis on diverse encoder architectures and datasets we highlight that even though SSL relies on large datasets and strong augmentations-both known in supervised learning as regularization techniques that reduce overfitting-still significant fractions of training data points experience high memorization. Through our empirical results, we show that this memorization is essential for encoders to achieve higher generalization performance on different downstream tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion ModelsDominik Hintersdorf, Lukas Struppek, Kristian Kersting, Adam Dziedzic 等NeurIPS 2024 · 被引用 46 次
- Evaluations of Machine Learning Privacy Defenses are MisleadingMichael Aerni, Jie Zhang, Florian TramèrCCS 2024 · 被引用 12 次
- Localizing Memorization in SSL Vision EncodersWenhao Wang, Adam Dziedzic, Michael Backes, Franziska BoenischNeurIPS 2024 · 被引用 11 次
- Exploring Structural Degradation in Dense Representations for Self-supervised LearningSiran Dai, Qianqian Xu, Peisong Wen, Yang Liu 等NeurIPS 2025 · 被引用 5 次
- Memorization in Graph Neural NetworksAdarsh Jamadandi, Jing Xu, Adam Dziedzic, Franziska BoenischNeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper24
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
相关 Paper
- Do SSL Models Have Déjà Vu? A Case of Unintended Memorization in Self-supervised LearningCasey Meehan, Florian Bordes, Pascal Vincent, Kamalika Chaudhuri 等NeurIPS 2023 · 被引用 26 次
- RAEncoder: A Label-Free Reversible Adversarial Examples Encoder for Dataset Intellectual Property ProtectionFan Xing, Zhuo Tian, Xuefeng Fan, Xiaoyi ZhouCVPR 2025
- SSLGuard: A Watermarking Scheme for Self-supervised Learning Pre-trained EncodersTianshuo Cong, Xinlei He, Yang ZhangCCS 2022 · 被引用 28 次
- Dataset Inference for Self-Supervised ModelsAdam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan 等NeurIPS 2022 · 被引用 59 次
- Reverse Engineering Self-Supervised LearningIdo Ben-Shaul, Ravid Shwartz-Ziv, Tomer Galanti, Shai Dekel 等NeurIPS 2023 · 被引用 55 次
