Mask-based Membership Inference Attacks for Retrieval-Augmented Generation
Mingrui Liu, Sixiao Zhang, Cheng Long
摘要
Retrieval-Augmented Generation (RAG) has been an effective approach to mitigate hallucinations in large language models (LLMs) by incorporating up-to-date and domain-specific knowledge. Recently, there has been a trend of storing up-to-date or copyrighted data in RAG knowledge databases instead of using it for LLM training. This practice has raised concerns about Membership Inference Attacks (MIAs), which aim to detect if a specific target document is stored in the RAG system's knowledge database so as to protect the rights of data producers. While research has focused on enhancing the trustworthiness of RAG systems, existing MIAs for RAG systems remain largely insufficient. Previous work either relies solely on the RAG system's judgment or is easily influenced by other documents or the LLM's internal knowledge, which is unreliable and lacks explainability. To address these limitations, we propose a Mask-Based Membership Inference Attacks (MBA) framework. Our framework first employs a masking algorithm that effectively masks a certain number of words in the target document. The masked text is then used to prompt the RAG system, and the RAG system is required to predict the mask values. If the target document appears in the knowledge database, the masked text will retrieve the complete target document as context, allowing for accurate mask prediction. Finally, we adopt a simple yet effective threshold-based method to infer the membership of target document by analyzing the accuracy of mask prediction. Our mask-based approach is more documentspecific, making the RAG system's generation less susceptible to distractions from other documents or the LLM's internal knowledge. Extensive experiments demonstrate the effectiveness of our approach compared to existing baseline models. CCS Concepts • Computing methodologies → Information extraction; • Security and privacy → Human and societal aspects of security and privacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image GenerationZiyuan Luo, Yangyi Zhao, Ka Chun Cheung, Simon See 等NeurIPS 2025 · 被引用 5 次
- Five Queries Are Enough: Query-Efficient and Surrogate-Free Membership Inference Attacks on RAG via EntailmentNguyen Linh Bao Nguyen, Wanlun Ma, Viet Vo, Alsharif Abuadbba 等USENIX Security 2026 · 被引用 4 次
- MrM: Black-Box Membership Inference Attacks Against Multimodal RAG SystemsPeiru Yang, Jinhua Yin, Haoran Zheng, Xueying Bai 等AAAI 2026 · 被引用 3 次
- Connect the Dots: Knowledge Graph–Guided Crawler Attack on Retrieval-Augmented Generation SystemsMengyu Yao, Ziqi Zhang, Ning Luo, Shaofei Li 等USENIX Security 2026 · 被引用 3 次
- RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache ReuseMingrui Liu, Sixiao Zhang, Cheng Long, Kwok Yan LamICML 2026 · 被引用 2 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
相关 Paper
- Riddle Me This! Stealthy Membership Inference for Retrieval-Augmented GenerationAli Naseh, Yuefeng Peng, Anshuman Suri, Harsh Chaudhari 等CCS 2025
- DCMI: A Differential Calibration Membership Inference Attack Against Retrieval-Augmented GenerationXinyu Gao, Xiangtao Meng, Yingkai Dong, Zheng Li 等CCS 2025
- Open Schrödinger's Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model ServicesYukun Jiang, Xinyue Shen, Michael Backes, Zheng Li 等ACL 2026
- WARP: A Word-Level Backdoor Attack Targeting RAG Systems via Retrieval Corpus PoisoningHui Liu, Yibo Zhou, Liguo Dong, Weidong Li 等KDD 2026
- RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language ModelsPeizhuo Lv, Mengjie Sun, Hao Wang, XiaoFeng Wang 等CCS 2025
