Governing Open Vocabulary Data Leaks Using an Edge LLM through Programming by Example
Qiyu Li, Jinhe Wen, Haojian Jin
摘要
A major concern with integrating large language model (LLM) services (e.g., ChatGPT) into workplaces is that employees may inadvertently leak sensitive information through their prompts. Since user prompts can involve arbitrary vocabularies, conventional data leak mitigation solutions, such as string-matching-based filtering, often fall short. We present GPTWall, a privacy firewall that helps internal admins create and manage policies to mitigate data leaks in prompts sent to external LLM services. GPTWall's key innovations are (1) introducing a lightweight LLM running on the edge to obfuscate target information in prompts and restore the information after receiving responses, and (2) helping admins author fine-grained disclosure policies through programming by example. We evaluated GPTWall with 12 participants and found that they could create an average of 17.7 policies within 30 minutes, achieving an increase of 29% in precision and 22% in recall over the state-of-the-art data de-identification tool.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Anti-adversarial Learning: Desensitizing Prompts for Large Language ModelXuan Li, Zhe Yin, Xiaodong Gu, Beijun ShenAAAI 2026
- Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM InferenceJiangou Zhan, Wenhui Zhang, Zheng Zhang, Huanran Xue 等AAAI 2025 · 被引用 10 次
- Operationalizing Data Minimization for Privacy-Preserving LLM PromptingJijie Zhou, Niloofar Mireshghallah, Tianshi LiICLR 2026 · 被引用 13 次
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina 等CCS 2024 · 被引用 28 次
- Instruction Backdoor Attacks Against Customized LLMsRui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang 等USENIX Security 2024 · 被引用 83 次
