Déjà Vu Memorization in Vision-Language Models
Bargav Jayaraman, Chuan Guo, Kamalika Chaudhuri
Abstract
Vision-Language Models (VLMs) have emerged as the state-of-the-art representation learning solution, with myriads of downstream applications such as image classification, retrieval and generation. A natural question is whether these models memorize their training data, which also has implications for generalization. We propose a new method for measuring memorization in VLMs, which we call déjà vu memorization. For VLMs trained on image-caption pairs, we show that the model indeed retains information about individual objects in the training images beyond what can be inferred from correlations or the image caption. We evaluate déjà vu memorization at both sample and population level, and show that it is significant for OpenCLIP trained on as many as 50M image-caption pairs. Finally, we show that text randomization considerably mitigates memorization while only moderately impacting the model's downstream task performance.
Prior work has looked into this problem for image-only representation models [Meehan et al., 2023] by measuring whether the model can predict the foreground of an image (e.g, black swan) beyond simple correlations based simply on its background (e.g, water). However, such simple solutions do not apply here. VLMs have two separate modalities -text and image, and the data sets used to train and evaluate them are considerably more complex than the simple foreground-background structure of ImageNet (see Figure 6 for an example). A consequence is that the image and text modalities 38th Conference on Neural Information Processing Systems (NeurIPS 2024).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 617529f6-66e7-4680-b8d5-f0fb6b8075c4Cited by top-tier papers7
- Measuring Dejavu Memorization EfficientlyNarine Kokhlikyan, Bargav Jayaraman, Florian Bordes, Chuan Guo et al.NeurIPS 2024 · 4 citations
- Quantifying Cross-Modality Memorization in Vision-Language ModelsYuxin Wen, Yangsibo Huang, Tom Goldstein, Ravi Kumar et al.NeurIPS 2025 · 3 citations
- FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User DataWenhao Wang, Zijie Yu, Rui Ye, Jianqing Zhang et al.EMNLP 2025
- Rethinking Membership Inference Attacks for CLIPLluís GómezAAAI 2026
- Captured by Captions: On Memorization and its Mitigation in CLIP ModelsWenhao Wang, Adam Dziedzic, Grace C. Kim, Michael Backes et al.ICLR 2025
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
Related papers
- Extracting Training Data From Document-Based VQA ModelsFrancesco Pinto, Nathalie Rauschmayr, Florian Tramèr, Philip Torr et al.ICML 2024 · 7 citations
- Analyzing and Mitigating Object Hallucination: A Training Bias PerspectiveYifan Li, Kun Zhou, Xin Zhao, Lei Fang et al.AAAI 2026 · 8 citations
- Teaching CLIP to Count to TenRoni Paiss, Ariel Ephrat, Omer Tov, Shiran Zada et al.ICCV 2023 · 196 citations
- Leveraging Vision-Language Models for Improving Domain Generalization in Image ClassificationSravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, R. Venkatesh BabuCVPR 2024
- C-CLIP: Multimodal Continual Learning for Vision-Language ModelWenzhuo Liu, Fei Zhu, Longhui Wei, Qi TianICLR 2025
