PLeak: Prompt Leaking Attacks against Large Language Model Applications
Bo Hui, Haolin Yuan, Neil Gong, Philippe Burlina, Yinzhi Cao
摘要
Large Language Models (LLMs) enable a new ecosystem with many downstream applications, called LLM applications, with different natural language processing tasks. The functionality and performance of an LLM application highly depend on its system prompt, which instructs the backend LLM on what task to perform. Therefore, an LLM application developer often keeps a system prompt confidential to protect its intellectual property. As a result, a natural attack, called prompt leaking, is to steal the system prompt from an LLM application, which compromises the developer's intellectual property. Existing prompt leaking attacks primarily rely on manually crafted queries, and thus achieve limited effectiveness. In this paper, we design a novel, closed-box prompt leaking attack framework, called PLeak, to optimize an adversarial query such that when the attacker sends it to a target LLM application, its response reveals its own system prompt. We formulate finding such an adversarial query as an optimization problem and solve it with a gradient-based method approximately. Our key idea is to break down the optimization goal by optimizing adversary queries for system prompts incrementally, i.e., starting from the first few tokens of each system prompt step by step until the entire length of the system prompt. We evaluate PLeak in both offline settings and for real-world LLM applications, e.g., those on Poe, a popular platform hosting such applications. Our results show that PLeak can effectively leak system prompts and significantly outperforms not only baselines that manually curate queries but also baselines with optimized queries that are modified and adapted from existing jailbreaking attacks. We responsibly reported the issues to Poe and are still waiting for their response. Our implementation is available at this repository: https://github.com/BHui97/PLeak .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Prompt Injection Attack to Tool Selection in LLM AgentsJiawen Shi, Zenghui Yuan, Guiyao Tie, Pan Zhou 等NDSS 2026 · 被引用 181 次
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan 等VLDB 2024 · 被引用 66 次
- PromptLocate: Localizing Prompt Injection AttacksYuqi Jia, Yupei Liu, Zedian Shao, Jinyuan Jia 等S&P 2026 · 被引用 35 次
- Optimization-based Prompt Injection Attack to LLM-as-a-JudgeJiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang 等CCS 2024 · 被引用 33 次
- ASIDE: Architectural Separation of Instructions and Data in Language ModelsEgor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova 等ICLR 2026 · 被引用 28 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
相关 Paper
- PRSA: Prompt Stealing Attacks against Real-World Prompt ServicesYong Yang, Changjiang Li, Qingming Li, Oubo Ma 等USENIX Security 2025
- Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based ApplicationsYong Yang, Chong Fu, Tong Zhang, Rui Zeng 等CCS 2026
- You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System VectorsBochuan Cao, Changjiang Li, Yuanpu Cao, Yameng Ge 等CCS 2025
- RepeatLeakage: Leak Prompts from Repeating as Large Language Model Is a Good RepeaterYu Peng, Lijie Zhang, Peizhuo Lv, Kai ChenAAAI 2025 · 被引用 3 次
- LLMThief: Evaluating Configuration Leaking Risks in Commercial LLM App StoresPinji Chen, Jinlong Jiang, Jianjun Chen, Feiran Qin 等S&P 2026 · 被引用 1 次
