Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based Applications
Yong Yang, Chong Fu, Tong Zhang, Rui Zeng, Qingming Li, Tianyu Du, Zonghui Wang, Shouling Ji, Wenzhi Chen
Abstract
Large language model (LLM)-based applications rely on system prompts to encode their core logic and developer-defined constraints, making them a critical form of intellectual property. However, these prompts are highly vulnerable to prompt leaking attacks. While the feasibility of such attacks has been demonstrated in controlled settings, a significant gap exists in understanding their prevalence, underlying mechanisms, and practical defenses within real-world deployments. In this paper, we bridge this gap by providing a systematic investigation into the landscape of prompt leaking in real-world LLMbased applications. Our study unfolds in three key aspects. First, we conduct a large-scale measurement of 1,200 applications across six major commercial platforms, revealing that over 80% of deployments leak system prompts under realistic adversarial queries, often exposing sensitive information like third-party API keys. Besides, our evaluation of existing defenses shows that they fail to prevent leakage without degrading usability. Second, to understand the root cause of these failures, we perform an attention-level mechanistic analysis and uncover a fundamental phenomenon we term attention drift, where query-key alignment bias and softmax amplification cause the LLMs to progressively ignore defensive constraints. Finally, guided by these insights, we propose AREA, a practical defense that re-anchors the LLM's attention via an optimizable soft prompt. Extensive experiments and real-world case studies demonstrate that AREA matches the leakage resistance of state-of-the-art defenses while improving average usability by over 33% and reducing optimization overhead by nearly 3×. The real-world significance of our work is further underscored by our responsible disclosure to affected vendors, two of whom have officially classified these leaks as medium-severity vulnerabilities. CCS Concepts • Security and privacy → Software security engineering; • Computing methodologies → Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff et al.USENIX Security 2026 · 134 citations
- Optimization-based Prompt Injection Attack to LLM-as-a-JudgeJiawen Shi, Zenghui Yuan, Yinuo Liu, Yue Huang et al.CCS 2024 · 33 citations
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
Related papers
- PRSA: Prompt Stealing Attacks against Real-World Prompt ServicesYong Yang, Changjiang Li, Qingming Li, Oubo Ma et al.USENIX Security 2025
- You Can't Steal Nothing: Mitigating Prompt Leakages in LLMs via System VectorsBochuan Cao, Changjiang Li, Yuanpu Cao, Yameng Ge et al.CCS 2025
- SecAlign: Defending Against Prompt Injection with Preference OptimizationSizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri et al.CCS 2025 · 1 citation
- Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM SecurityXiang Fang, Wanlong FangAAAI 2026 · 4 citations
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li et al.ACL 2026 · 8 citations
