Understanding and Detecting File Knowledge Leakage in GPT App Ecosystem
Chuan Yan, Bowei Guan, Yazhi Li, Mark Huasong Meng, Liuhuo Wan, Guangdong Bai
摘要
OpenAI has enabled third-party developers to build applications around ChatGPT, known as GPTs, to expand its capability to handle complex and specialized tasks. A key feature of GPTs is Retrieval-Augmented Generation (RAG), which allows developers to upload documents containing domain knowledge or application context, referred to as file knowledge. However, these documents often contain sensitive information, and the security mechanisms governing access control in GPTs remains an underexplored area. In this work, we present the first comprehensive study on file knowledge leakage within GPTs. We develop GPTs-Filtor, leveraging the unique characteristics of GPTs deployment, to perform an in-depth analysis and detection of file knowledge leakage at both user interaction (i.e., prompt) and network transmission levels. Applying GPTs-Filtor to 8,000 popular GPTs across eight different categories, we reveal widespread vulnerabilities in the current GPTs development and deployment model. We detect 618 cases of leakage among 1,331 GPTs that involve uploaded file knowledge, leading to the exfiltration of 3,645 file contents that contain highly-sensitive data such as internal bank audit transaction records. Our work underscores the pressing need for improved security practices in GPTs development and deployment, providing crucial insights for the secure development of this young but rapidly evolving ecosystem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLMThief: Evaluating Configuration Leaking Risks in Commercial LLM App StoresPinji Chen, Jinlong Jiang, Jianjun Chen, Feiran Qin 等S&P 2026 · 被引用 1 次
- PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning DetectionJungmin Lee, Peizhuo Lv, Yeonjoon LeeACL 2026
它引用的顶会 Paper9
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui 等ICLR 2024 · 被引用 146 次
- Optimization-based Prompt Injection Attack to LLM-as-a-JudgeJiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang 等CCS 2024 · 被引用 33 次
- CORELOCKER: Neuron-level Usage ControlZihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun 等S&P 2024 · 被引用 11 次
相关 Paper
- When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTsXinyue Shen, Yun Shen, Michael Backes, Yang ZhangACL 2025
- GPTracker: A Large-Scale Measurement of Misused GPTsXinyue Shen, Yun Shen, Michael Backes, Yang ZhangS&P 2025
- Exploring ChatGPT App Ecosystem: Distribution, Deployment and SecurityChuan Yan, Ruomai Ren, Mark Huasong Meng, Liuhuo Wan 等ASE 2024 · 被引用 3 次
- On the Vulnerability of Applying Retrieval-Augmented Generation within Knowledge-Intensive Application DomainsXun Xian, Ganghua Wang, Xuan Bi, Rui Zhang 等ICML 2025
- RepeatLeakage: Leak Prompts from Repeating as Large Language Model Is a Good RepeaterYu Peng, Lijie Zhang, Peizhuo Lv, Kai ChenAAAI 2025 · 被引用 3 次
