USENIX Security2026Top-tier venue
The Prompt Stealing Fallacy: Rethinking Metrics, Attacks, and Defenses
Zehang Deng, Haoyang Li, Wanlun Ma, Ruoxi Sun, Derui Wang, Minhui Xue, Haibo Hu, Sheng Wen, Yang Xiang
Abstract
Text-to-image (T2I) models are increasingly embedded in creative workflows, where well-crafted prompts function as valuable forms of intellectual property (IP). However, these models are susceptible to prompt stealing attacks (PSAs), where adversaries aim to reconstruct the original prompts used to generate images. In this paper, 1) we identify key shortcomings in current evaluation practices and propose two improved metrics: Style Similarity (SS) and a novel Prompt Significance (PS) score, which together provide a more faithful assessment of PSA effectiveness. Rather than existing metrics that rely solely on semantic similarity between original and stolen information across text or image modalities, the new metrics PS and SS assess attack effectiveness with a more practical focus by explicitly accounting for the importance of modifiers and the style replication of images generated from stolen prompts. 2) Through extensive evaluation using these metrics, we find that existing PSA methods, ranging from soft prompt stealing in white-box settings to hard prompt stealing in black-box settings, are not as effective as reported, especially in recovering high-contribution prompt components. We attribute this to fundamental constraints: white-box methods suffer from mismatched optimization objectives that poorly align with token-level visual semantics, while black-box approaches experience severe information loss due to their decoupling from the target T2I model's generation process. 3) We further introduce PromptThief, a black-box PSA framework that addresses the information loss in prior methods by leveraging reinforcement learning with Semantic Text-Text Similarity (STS) and SS to guide high token-level contribution recovery. PromptThief significantly outperforms existing baselines across multiple metrics and real-world scenarios. 4) We propose and evaluate two defense mechanisms: an adversarialexample-based active approach and a passive scheme through feature-level prompt watermarking. Our evaluation reveals that the active defense offers only limited robustness against adaptive PSAs, highlighting the need for further exploration * Corresponding author. in this direction. In contrast, the passive watermarking scheme demonstrates strong and consistent detection performance, even under various image transformations, offering a practical and reliable path forward for prompt IP protection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum et al.NeurIPS 2023 · 454 citations
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 304 citations
- Optimizing Prompts for Text-to-Image GenerationYaru Hao, Zewen Chi, Li Dong, Furu WeiNeurIPS 2023 · 303 citations
Related papers
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and VLM-Guided OptimizationMingzhe Li, Renhao 'Norman' Zhang, Zhiyang Wen, Siqi Pan et al.CVPR 2026
- Prompt Stealing Attacks Against Text-to-Image Generation ModelsXinyue Shen, Yiting Qu, Michael Backes, Yang ZhangUSENIX Security 2024 · 65 citations
- Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion ModelsShiqian Zhao, Chong Wang, Yiming Li, Yihao Huang et al.NDSS 2026 · 3 citations
- On the Effectiveness of Prompt Stealing Attacks on In-the-Wild PromptsYicong Tan, Xinyue Shen, Yun Shen, Michael Backes et al.S&P 2025
- Prompt Obfuscation for Large Language ModelsDavid Pape, Sina Mavali, Thorsten Eisenhofer, Lea SchönherrUSENIX Security 2025
