Déjà Vu Memorization in Vision-Language Models
Bargav Jayaraman, Chuan Guo, Kamalika Chaudhuri
摘要
Vision-Language Models (VLMs) have emerged as the state-of-the-art representation learning solution, with myriads of downstream applications such as image classification, retrieval and generation. A natural question is whether these models memorize their training data, which also has implications for generalization. We propose a new method for measuring memorization in VLMs, which we call déjà vu memorization. For VLMs trained on image-caption pairs, we show that the model indeed retains information about individual objects in the training images beyond what can be inferred from correlations or the image caption. We evaluate déjà vu memorization at both sample and population level, and show that it is significant for OpenCLIP trained on as many as 50M image-caption pairs. Finally, we show that text randomization considerably mitigates memorization while only moderately impacting the model's downstream task performance.
Prior work has looked into this problem for image-only representation models [Meehan et al., 2023] by measuring whether the model can predict the foreground of an image (e.g, black swan) beyond simple correlations based simply on its background (e.g, water). However, such simple solutions do not apply here. VLMs have two separate modalities -text and image, and the data sets used to train and evaluate them are considerably more complex than the simple foreground-background structure of ImageNet (see Figure 6 for an example). A consequence is that the image and text modalities 38th Conference on Neural Information Processing Systems (NeurIPS 2024).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Measuring Dejavu Memorization EfficientlyNarine Kokhlikyan, Bargav Jayaraman, Florian Bordes, Chuan Guo 等NeurIPS 2024 · 被引用 4 次
- Quantifying Cross-Modality Memorization in Vision-Language ModelsYuxin Wen, Yangsibo Huang, Tom Goldstein, Ravi Kumar 等NeurIPS 2025 · 被引用 3 次
- FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User DataWenhao Wang, Zijie Yu, Rui Ye, Jianqing Zhang 等EMNLP 2025
- Rethinking Membership Inference Attacks for CLIPLluís GómezAAAI 2026
- Captured by Captions: On Memorization and its Mitigation in CLIP ModelsWenhao Wang, Adam Dziedzic, Grace C. Kim, Michael Backes 等ICLR 2025
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
相关 Paper
- Extracting Training Data From Document-Based VQA ModelsFrancesco Pinto, Nathalie Rauschmayr, Florian Tramèr, Philip Torr 等ICML 2024 · 被引用 7 次
- Analyzing and Mitigating Object Hallucination: A Training Bias PerspectiveYifan Li, Kun Zhou, Xin Zhao, Lei Fang 等AAAI 2026 · 被引用 8 次
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada 等ICCV 2023 · 被引用 196 次
- Leveraging Vision-Language Models for Improving Domain Generalization in Image ClassificationSravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, R. Venkatesh BabuCVPR 2024
- C-CLIP: Multimodal Continual Learning for Vision-Language ModelWenzhuo Liu, Fei Zhu, Longhui Wei, Qi TianICLR 2025
