Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
Jiankun Zhang, Shenglai Zeng, Jie Ren, Tianqi Zheng, Hui Liu, Xianfeng Tang, Hui Liu, Yi Chang
Abstract
Multimodal Retrieval-Augmented Generation (MRAG) systems enhance LMMs by integrating external multimodal databases, but introduce unexplored privacy vulnerabilities.While text-based RAG privacy risks have been studied, multimodal data presents unique challenges.We provide the first systematic analysis of MRAG privacy vulnerabilities across vision-language and speech-language modalities.Using a novel compositional structured prompt attack in a black-box setting, we demonstrate how attackers can extract private information by manipulating queries.Our experiments reveal that LMMs can both directly generate outputs resembling retrieved content and produce descriptions that indirectly expose sensitive information, highlighting the urgent need for robust privacy-preserving MRAG techniques.The code is available at https://github.com/phycholosogy/MRAGprivacy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Privacy-Aware Decoding: Mitigating Privacy Leakage of Large Language Models in Retrieval-Augmented GenerationHaoran Wang, Xiongxiao Xu, Baixiang Huang, Kai ShuKDD 2026 · 13 citations
- Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented GenerationShenglai Zeng, Jiankun Zhang, Kai Guo, Xinnan Dai et al.ICML 2026 · 1 citation
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems by Exploiting Knowledge AsymmetryYufei Chen, Yao Wang, Haibin Zhang, Tao GuICLR 2026 · 2 citations
- MrM: Black-Box Membership Inference Attacks Against Multimodal RAG SystemsPeiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai et al.AAAI 2026 · 3 citations
- PoisonedEye: Knowledge Poisoning Attack on Retrieval-Augmented Generation based Large Vision-Language ModelsChenyang Zhang, Xiaoyu Zhang, Jian Lou, Kai Wu et al.ICML 2025
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation SystemsZhenting Qi, Hanlin Zhang, Eric P. Xing, Sham M. Kakade et al.ICLR 2025
