Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications
Yong Yang, Chong Fu, Tong Zhang, Rui Zeng, Qingming Li, Tianyu Du, Zonghui Wang, Shouling Ji, Wenzhi Chen
摘要
Large language model (LLM)-based applications rely on system prompts to encode their core logic and developer-defined constraints, making them a critical form of intellectual property. However, these prompts are highly vulnerable to prompt leaking attacks. While the feasibility of such attacks has been demonstrated in controlled settings, a significant gap exists in understanding their prevalence, underlying mechanisms, and practical defenses within real-world deployments. In this paper, we bridge this gap by providing a systematic investigation into the landscape of prompt leaking in real-world LLMbased applications. Our study unfolds in three key aspects. First, we conduct a large-scale measurement of 1,200 applications across six major commercial platforms, revealing that over 80% of deployments leak system prompts under realistic adversarial queries, often exposing sensitive information like third-party API keys. Besides, our evaluation of existing defenses shows that they fail to prevent leakage without degrading usability. Second, to understand the root cause of these failures, we perform an attention-level mechanistic analysis and uncover a fundamental phenomenon we term attention drift, where query-key alignment bias and softmax amplification cause the LLMs to progressively ignore defensive constraints. Finally, guided by these insights, we propose AREA, a practical defense that re-anchors the LLM's attention via an optimizable soft prompt. Extensive experiments and real-world case studies demonstrate that AREA matches the leakage resistance of state-of-the-art defenses while improving average usability by over 33% and reducing optimization overhead by nearly 3×. The real-world significance of our work is further underscored by our responsible disclosure to affected vendors, two of whom have officially classified these leaks as medium-severity vulnerabilities. CCS Concepts • Security and privacy → Software security engineering; • Computing methodologies → Natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia 等USENIX Security 2024 · 被引用 308 次
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff 等USENIX Security 2026 · 被引用 134 次
- Optimization-based Prompt Injection Attack to LLM-as-a-JudgeJiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang 等CCS 2024 · 被引用 33 次
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina 等CCS 2024 · 被引用 28 次
相关 Paper
- PRSA: Prompt Stealing Attacks against Real-World Prompt ServicesYong Yang, Changjiang Li, Qingming Li, Oubo Ma 等USENIX Security 2025
- You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System VectorsBochuan Cao, Changjiang Li, Yuanpu Cao, Yameng Ge 等CCS 2025
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri 等CCS 2025 · 被引用 1 次
- Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM SecurityXiang Fang, Wanlong FangAAAI 2026 · 被引用 4 次
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li 等ACL 2026 · 被引用 8 次
