Governing Open Vocabulary Data Leaks Using an Edge LLM through Programming by Example
Qiyu Li, Jinhe Wen, Haojian Jin
Abstract
A major concern with integrating large language model (LLM) services (e.g., ChatGPT) into workplaces is that employees may inadvertently leak sensitive information through their prompts. Since user prompts can involve arbitrary vocabularies, conventional data leak mitigation solutions, such as string-matching-based filtering, often fall short. We present GPTWall, a privacy firewall that helps internal admins create and manage policies to mitigate data leaks in prompts sent to external LLM services. GPTWall's key innovations are (1) introducing a lightweight LLM running on the edge to obfuscate target information in prompts and restore the information after receiving responses, and (2) helping admins author fine-grained disclosure policies through programming by example. We evaluated GPTWall with 12 participants and found that they could create an average of 17.7 policies within 30 minutes, achieving an increase of 29% in precision and 22% in recall over the state-of-the-art data de-identification tool.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Anti-adversarial Learning: Desensitizing Prompts for Large Language ModelXuan Li, Zhe Yin, Xiaodong Gu, Beijun ShenAAAI 2026
- Portcullis: A Scalable and Verifiable Privacy Gateway for Third-Party LLM InferenceJiangou Zhan, Wenhui Zhang, Zheng Zhang, Huanran Xue et al.AAAI 2025 · 10 citations
- Operationalizing Data Minimization for Privacy-Preserving LLM PromptingJijie Zhou, Niloofar Mireshghallah, Tianshi LiICLR 2026 · 13 citations
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
- Instruction Backdoor Attacks Against Customized LLMsRui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang et al.USENIX Security 2024 · 83 citations
