USENIX Security2025Top-tier venue
Prompt Obfuscation for Large Language Models
David Pape, Sina Mavali, Thorsten Eisenhofer, Lea Schönherr
Abstract
System prompts that include detailed instructions to describe the task performed by the underlying LLM can easily transform foundation models into tools and services with minimal overhead. They are often considered intellectual property, similar to the code of a software product, because of their crucial impact on the utility. However, extracting system prompts is easily possible. As of today, there is no effective countermeasure to prevent the stealing of system prompts, and all safeguarding efforts could be evaded. In this work, we propose an alternative to conventional system prompts. We introduce prompt obfuscation to prevent the extraction of the system prompt with little overhead. The core idea is to find a representation of the original system prompt that leads to the same functionality, while the obfuscated system prompt does not contain any information that allows conclusions to be drawn about the original system prompt. We evaluate our approach by comparing our obfuscated prompt output with the output of the original prompt, using eight distinct metrics to measure the lexical, character-level, and semantic similarity. We show that the obfuscated version is constantly on par with the original one. We further perform three different deobfuscation attacks with varying attacker knowledge--covering both black-box and white-box conditions--and show that in realistic attack scenarios an attacker is unable to extract meaningful information. Overall, we demonstrate that prompt obfuscation is an effective mechanism to safeguard the intellectual property of a system prompt while maintaining the same utility as the original prompt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu et al.EMNLP 2025 · 9 citations
- Leaky Thoughts: Large Reasoning Models Are Not Private ThinkersTommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun et al.EMNLP 2025 · 2 citations
- Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based ApplicationsYong Yang, Chong Fu, Tong Zhang, Rui Zeng et al.CCS 2026
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMsNafiseh Nikeghbal, Amir Hossein Kargaran, Jana DiesnerEMNLP 2025
- Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and VerificationTuc Nguyen, Yifan Hu, Thai LeEMNLP 2025
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum et al.NeurIPS 2023 · 454 citations
Related papers
- Uncovering Prompt Elements: Cloning System Prompts from Behavioral TracesYi Qian, Fei Peng, Hao Wu, Ligeng Chen et al.ASE 2025 · 1 citation
- PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable PromptsQinfeng Li, Yuntai Bao, Jianghui Hu, Wenqi Zhang et al.ICML 2026
- Web Intellectual Property at Risk: Preventing Unauthorized Real-Time Retrieval by Large Language ModelsYisheng Zhong, Yizhu Wen, Junfeng Guo, Mehran Kafai et al.EMNLP 2025
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
- Anti-adversarial Learning: Desensitizing Prompts for Large Language ModelXuan Li, Zhe Yin, Xiaodong Gu, Beijun ShenAAAI 2026
