USENIX Security2026Top-tier venue
Connect the Dots: Knowledge Graph–Guided Crawler Attack on Retrieval-Augmented Generation Systems
Mengyu Yao, Ziqi Zhang, Ning Luo, Shaofei Li, Yifeng Cai, Xiangqun Chen, Yao Guo, Ding Li
Abstract
Stealing attacks pose a persistent threat to the intellectual property of deployed machine-learning systems. Retrieval-augmented generation (RAG) intensifies this risk by extending the attack surface beyond model weights to knowledge base that often contains IP-bearing assets such as proprietary runbooks, curated domain collections, or licensed documents. Recent work shows that multi-turn questioning can gradually steal corpus content from RAG systems, yet existing attacks are largely heuristic and often plateau early. We address this gap by formulating RAG knowledge-base stealing as an adaptive stochastic coverage problem (ASCP), where each query is a stochastic action and the attacker's goal is to maximize the conditional expected marginal gain (CMG) in corpus coverage under a query budget. Bridging ASCP to real-world black-box RAG knowledge-base stealing raises three challenges: CMG is unobservable, the natural-language action space is intractably large, and feasibility constraints require stealthy queries that remain effective under diverse architectures. We introduce RAGCrawler, a knowledge graph-guided attacker that maintains a global attacker-side state to estimate coverage gains, schedule high-value semantic anchors, and generate non-redundant natural queries. Across four corpora and four generators with BGE retriever, RAGCrawler achieves 66.8% average coverage (up to 84.4%) within 1,000 queries, improving coverage by 44.90% relative to the strongest baseline. It also reduces the queries needed to reach 70% coverage by at least $4.03x on average and enables surrogate reconstruction with answer similarity up to 0.699. Moreover, the attack remains effective under retriever switching and newer RAG techniques such as query rewriting and multi-query retrieval, while being difficult for existing defenses to block. These results highlight urgent needs to protect RAG knowledge assets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a21ca557-dd58-492e-aa3d-f11aa5eabf31Builds on28
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- HOLMES: Real-Time APT Detection through Correlation of Suspicious Information FlowsSadegh Momeni Milajerdi, Rigel Gjomemo, Birhanu Eshete, R. Sekar et al.S&P 2019 · 550 citations
- Stealing Links from Graph Neural NetworksXinlei He, Jinyuan Jia, Michael Backes, Neil Zhenqiang Gong et al.USENIX Security 2021 · 226 citations
- Query Rewriting in Retrieval-Augmented Large Language ModelsXinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao et al.EMNLP 2023 · 191 citations
Related papers
- Query-Efficient Agentic Graph Extraction Attacks on GraphRAG SystemsShuhua Yang, Jiahao Zhang, Yilong Wang, Dongwon Lee et al.ACL 2026 · 2 citations
- RAG-WM: An Efficient Black-Box Watermarking Approach for Retrieval-Augmented Generation of Large Language ModelsPeizhuo Lv, Mengjie Sun, Hao Wang, XiaoFeng Wang et al.CCS 2025
- Towards Whole-corpus Reconstruction of Heterogeneous RAG Knowledge BasesPeiru Yang, Yi Luo, Zhenfeng Gao, Tong Ju et al.ICML 2026
- KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented GenerationQizhi Chen, Chao Qi, Yihong Huang, Muquan Li et al.WWW 2026
- RAGFort: Dual-Path Defense Against Proprietary Knowledge Base Extraction in Retrieval-Augmented GenerationQinfeng Li, Miao Pan, Ke Xiong, Ge Su et al.AAAI 2026
