Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
Jiankun Zhang, Shenglai Zeng, Jie Ren, Tianqi Zheng, Hui Liu, Xianfeng Tang, Hui Liu, Yi Chang
摘要
Multimodal Retrieval-Augmented Generation (MRAG) systems enhance LMMs by integrating external multimodal databases, but introduce unexplored privacy vulnerabilities.While text-based RAG privacy risks have been studied, multimodal data presents unique challenges.We provide the first systematic analysis of MRAG privacy vulnerabilities across vision-language and speech-language modalities.Using a novel compositional structured prompt attack in a black-box setting, we demonstrate how attackers can extract private information by manipulating queries.Our experiments reveal that LMMs can both directly generate outputs resembling retrieved content and produce descriptions that indirectly expose sensitive information, highlighting the urgent need for robust privacy-preserving MRAG techniques.The code is available at https://github.com/phycholosogy/MRAGprivacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Privacy-Aware Decoding: Mitigating Privacy Leakage of Large Language Models in Retrieval-Augmented GenerationHaoran Wang, Xiongxiao Xu, Baixiang Huang, Kai ShuKDD 2026 · 被引用 13 次
- Fix Before Search: Benchmarking Agentic Visual Query Pre-processing in Multimodal Retrieval-augmented GenerationShenglai Zeng, Jiankun Zhang, Kai Guo, Xinnan Dai 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems by Exploiting Knowledge AsymmetryYufei Chen, Yao Wang, Haibin Zhang, Tao GuICLR 2026 · 被引用 2 次
- MrM: Black-Box Membership Inference Attacks Against Multimodal RAG SystemsPeiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai 等AAAI 2026 · 被引用 3 次
- PoisonedEye: Knowledge Poisoning Attack on Retrieval-Augmented Generation based Large Vision-Language ModelsChenyang Zhang, Xiaoyu Zhang, Jian Lou, Kai Wu 等ICML 2025
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 被引用 6 次
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation SystemsZhenting Qi, Hanlin Zhang, Eric P. Xing, Sham M. Kakade 等ICLR 2025
