Localizing Memorization in SSL Vision Encoders
Wenhao Wang, Adam Dziedzic, Michael Backes, Franziska Boenisch
摘要
Recent work on studying memorization in self-supervised learning (SSL) suggests that even though SSL encoders are trained on millions of images, they still memorize individual data points. While effort has been put into characterizing the memorized data and linking encoder memorization to downstream utility, little is known about where the memorization happens inside SSL encoders. To close this gap, we propose two metrics for localizing memorization in SSL encoders on a per-layer (layermem) and per-unit basis (unitmem). Our localization methods are independent of the downstream task, do not require any label information, and can be performed in a forward pass. By localizing memorization in various encoder architectures (convolutional and transformer-based) trained on diverse datasets with contrastive and non-contrastive SSL frameworks, we find that (1) while SSL memorization increases with layer depth, highly memorizing units are distributed across the entire encoder, (2) a significant fraction of units in SSL encoders experiences surprisingly high memorization of individual data points, which is in contrast to models trained under supervision, (3) atypical (or outlier) data points cause much higher layer and unit memorization than standard data points, and (4) in vision transformers, most memorization happens in the fully-connected layers. Finally, we show that localizing memorization in SSL has the potential to improve fine-tuning and to inform pruning strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Exploring Structural Degradation in Dense Representations for Self-supervised LearningSiran Dai, Qianqian Xu, Peisong Wen, Yang Liu 等NeurIPS 2025 · 被引用 5 次
- Demystifying Foreground-Background Memorization in Diffusion ModelsJimmy Z. Di, Yiwei Lu, Yaoliang Yu, Gautam Kamath 等AAAI 2026 · 被引用 1 次
- Privacy Attacks on Image AutoRegressive ModelsAntoni Kowalczuk, Jan Dubinski, Franziska Boenisch, Adam DziedzicICML 2025
- Captured by Captions: On Memorization and its Mitigation in CLIP ModelsWenhao Wang, Adam Dziedzic, Grace C. Kim, Michael Backes 等ICLR 2025
它引用的顶会 Paper27
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
相关 Paper
- Memorization in Self-Supervised Learning Improves Downstream GeneralizationWenhao Wang, Muhammad Ahmad Kaleem, Adam Dziedzic, Michael Backes 等ICLR 2024 · 被引用 19 次
- Three Guidelines You Should Know for Universally Slimmable Self-Supervised LearningYun-Hao Cao, Peiqin Sun, Shuchang ZhouCVPR 2023
- Do SSL Models Have Déjà Vu? A Case of Unintended Memorization in Self-supervised LearningCasey Meehan, Florian Bordes, Pascal Vincent, Kamalika Chaudhuri 等NeurIPS 2023 · 被引用 26 次
- Rethinking Federated Unlearning via the Lens of MemorizationJiaheng Wei, Yanjun Zhang, He Zhang, Leo Yu Zhang 等KDD 2026 · 被引用 1 次
- In-Context Symmetries: Self-Supervised Learning through Contextual World ModelsSharut Gupta, Chenyu Wang, Yifei Wang, Tommi S. Jaakkola 等NeurIPS 2024 · 被引用 8 次
