USENIX Security2026Top-tier venue
Bypassing Prompt Guards in Production with Controlled-Release Prompting
Jaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan Wang
Abstract
Ball et al. [5] recently established that prompt filtering for AI alignment faces a fundamental barrier: under standard cryptographic assumptions, no filter running significantly faster than the protected model can universally distinguish adversarial prompts from benign ones. We investigate whether this impossibility result translates to real-world vulnerabilities in deployed large language model (LLM) systems. We answer affirmatively by introducing controlled-release prompting, a practical instantiation of the theoretical framework that exploits the resource asymmetry between lightweight input filters and the main models they protect. Unlike the theoretical construction, our attack does not require model modification: it generates malicious prompts that are indecipherable by any bounded filter yet remain tractable to the target LLM. We find our attack to be successful on four major chat platforms (Google Gemini, DeepSeek Chat, xAI Grok, and Mistral Le Chat) where baseline methods fail. Additionally, we apply our attack to extract copyrighted data from Gemini. Finally, we provide a systematic evaluation of 14 open-weight prompt guard models, revealing that even reasoning-capable filters cannot reliably detect our attack without incurring prohibitive resource overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 082d8ac7-ad32-47d2-a51c-63de8d53ec35Builds on18
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma et al.NeurIPS 2024 · 338 citations
- COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityXingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin et al.ICML 2024 · 173 citations
Related papers
- On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI AlignmentSarah Ball, Greg Gluch, Shafi Goldwasser, Frauke Kreuter et al.ICLR 2026 · 16 citations
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-Based Prompt Injection Attacks via the Fine-Tuning InterfaceAndrey Labunets, Nishit V. Pandya, Ashish Hooda, Xiaohan Fu et al.S&P 2025
- Scalable Extraction of Training Data from Aligned, Production Language ModelsMilad Nasr, Javier Rando, Nicholas Carlini, Jonathan Hayase et al.ICLR 2025
- Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking CompetitionSander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard et al.EMNLP 2023 · 25 citations
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
