PLeak: Prompt Leaking Attacks against Large Language Model Applications
Bo Hui, Haolin Yuan, Neil Gong, Philippe Burlina, Yinzhi Cao
Abstract
Large Language Models (LLMs) enable a new ecosystem with many downstream applications, called LLM applications, with different natural language processing tasks. The functionality and performance of an LLM application highly depend on its system prompt, which instructs the backend LLM on what task to perform. Therefore, an LLM application developer often keeps a system prompt confidential to protect its intellectual property. As a result, a natural attack, called prompt leaking, is to steal the system prompt from an LLM application, which compromises the developer's intellectual property. Existing prompt leaking attacks primarily rely on manually crafted queries, and thus achieve limited effectiveness. In this paper, we design a novel, closed-box prompt leaking attack framework, called PLeak, to optimize an adversarial query such that when the attacker sends it to a target LLM application, its response reveals its own system prompt. We formulate finding such an adversarial query as an optimization problem and solve it with a gradient-based method approximately. Our key idea is to break down the optimization goal by optimizing adversary queries for system prompts incrementally, i.e., starting from the first few tokens of each system prompt step by step until the entire length of the system prompt. We evaluate PLeak in both offline settings and for real-world LLM applications, e.g., those on Poe, a popular platform hosting such applications. Our results show that PLeak can effectively leak system prompts and significantly outperforms not only baselines that manually curate queries but also baselines with optimized queries that are modified and adapted from existing jailbreaking attacks. We responsibly reported the issues to Poe and are still waiting for their response. Our implementation is available at this repository: https://github.com/BHui97/PLeak .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8fd8ad1-a43f-4f3f-a80c-fca6bb57d4b6Cited by top-tier papers52
- Prompt Injection Attack to Tool Selection in LLM AgentsJiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou et al.NDSS 2026 · 181 citations
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan et al.VLDB 2024 · 66 citations
- PromptLocate: Localizing Prompt Injection AttacksYuqi Jia, Yupei Liu, Zedian Shao, Jinyuan Jia et al.S&P 2026 · 35 citations
- Optimization-based Prompt Injection Attack to LLM-as-a-JudgeJiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang et al.CCS 2024 · 33 citations
- ASIDE: Architectural Separation of Instructions and Data in Language ModelsEgor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova et al.ICLR 2026 · 28 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
Related papers
- PRSA: Prompt Stealing Attacks against Real-World Prompt ServicesYong Yang, Changjiang Li, Qingming Li, Oubo Ma et al.USENIX Security 2025
- Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based ApplicationsYong Yang, Chong Fu, Tong Zhang, Rui Zeng et al.CCS 2026
- You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System VectorsBochuan Cao, Changjiang Li, Yuanpu Cao, Yameng Ge et al.CCS 2025
- RepeatLeakage: Leak Prompts from Repeating as Large Language Model Is a Good RepeaterYu Peng, Lijie Zhang, Peizhuo Lv, Kai ChenAAAI 2025 · 3 citations
- LLMThief: Evaluating Configuration Leaking Risks in Commercial LLM App StoresPinji Chen, Jinlong Jiang, Jianjun Chen, Feiran Qin et al.S&P 2026 · 1 citation
