RAGFort: Dual-Path Defense Against Proprietary Knowledge Base Extraction in Retrieval-Augmented Generation
Qinfeng Li, Miao Pan, Ke Xiong, Ge Su, Zhiqiang Shen, Yan Liu, Bing Sun, Hao Peng, Xuhong Zhang
摘要
Retrieval-Augmented Generation (RAG) systems deployed over proprietary knowledge bases face growing threats from reconstruction attacks that aggregate model responses to replicate knowledge bases. Such attacks exploit both intra-class and inter-class paths—progressively extracting fine-grained knowledge within topics and diffusing it across semantically related ones, thereby enabling comprehensive extraction of the original knowledge base. However, existing defenses target only one path, leaving the other unprotected. We conduct a systematic exploration to assess the impact of protecting each path independently and find that joint protection is essential for effective defense. Based on this, we propose RAGFort, a structure-aware dual-module defense combining contrastive reindexing for inter-class isolation and constrained cascade generation for intra-class protection. Experiments across security, performance, and robustness confirm that RAGFort significantly reduces reconstruction success while preserving answer quality, offering the first comprehensive defense against knowledge base extraction attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- DistillSpec: Improving Speculative Decoding via Knowledge DistillationYongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 143 次
- Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic DataShenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren 等EMNLP 2025 · 被引用 7 次
- Faster Cascades via Speculative DecodingHarikrishna Narasimhan, Wittawat Jitkrittum, Ankit Singh Rawat, Seungyeon Kim 等ICLR 2025
相关 Paper
- Safe RAG by RAG: Untying the Bell That RAG Rang with the RAG HandXun Liang, Mengwei Wang, Yuefeng Ma, Simin NiuAAAI 2026
- SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAGXiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu 等NeurIPS 2025 · 被引用 16 次
- IRAG: Robust Multimodal Retrieval-Augmented Generation via Hazard SeparationRuikun Luo, Zixiao Feng, Lin Gu, Xiaoyu XiaWWW 2026
- Towards Whole-corpus Reconstruction of Heterogeneous RAG Knowledge BasesPeiru Yang, Yi Luo, Zhenfeng Gao, Tong Ju 等ICML 2026
- Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems by Exploiting Knowledge AsymmetryYufei Chen, Yao Wang, Haibin Zhang, Tao GuICLR 2026 · 被引用 2 次
