Extracting Training Data From Document-Based VQA Models
Francesco Pinto, Nathalie Rauschmayr, Florian Tramèr, Philip Torr, Federico Tombari
Abstract
Vision-Language Models (VLMs) have made remarkable progress in document-based Visual Question Answering (i.e., responding to queries about the contents of an input document provided as an image). In this work, we show these models can memorize responses for training samples and regurgitate them even when the relevant visual information has been removed. This includes Personal Identifiable Information (PII) repeated once in the training set, indicating these models could divulge memorised sensitive information and therefore pose a privacy risk. We quantitatively measure the extractability of information in controlled experiments and differentiate between cases where it arises from generalization capabilities or from memorization. We further investigate the factors that influence memorization across multiple state-of-the-art models and propose an effective heuristic countermeasure that empirically prevents the extractability of PII.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- TokenSwap: A Lightweight Method to Disrupt Memorized Sequences in LLMsParjanya Prajakta Prashant, Kaustubh Ponkshe, Babak SalimiNeurIPS 2025 · 2 citations
- Copyright-Protected Language Generation via Adaptive Model FusionJavier Abad, Konstantin Donhauser, Francesco Pinto, Fanny YangICLR 2025
- MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation ModelsChejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie et al.ICLR 2025
- DocMIA: Document-Level Membership Inference Attacks against DocVQA ModelsKhanh Nguyen, Raouf Kerkouche, Mario Fritz, Dimosthenis KaratzasICLR 2025
- Black-Box Membership Inference Attacks for Video Training Data in Multimodal Large Language ModelsJinrui Wang, Zhenfeng Gao, Wendan Wang, Huili Wang et al.ACL 2026
Builds on19
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
Related papers
- Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language ModelsElena Sofia Ruzzetti, Giancarlo A. Xompero, Davide Venditti, Fabio Massimo ZanzottoACL 2025 · 9 citations
- Rethinking the Role of Verbatim Memorization in LLM PrivacyTom Sander, Bargav Jayaraman, Mark Ibrahim, Kamalika Chaudhuri et al.NeurIPS 2025 · 5 citations
- Exploiting the Shadows: Unveiling Privacy Leaks through Lower-Ranked Tokens in Large Language ModelsYuan Zhou, Zhuo Zhang, Xiangyu ZhangACL 2025 · 2 citations
- Private Attribute Inference from Images with Vision-Language ModelsBatuhan Tömekçe, Mark Vero, Robin Staab, Martin T. VechevNeurIPS 2024 · 54 citations
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language ModelsXiongtao Sun, HUI LI, Jiaming Zhang, Yujie Yang et al.ICML 2026 · 3 citations
