Uncovering Prompt Elements: Cloning System Prompts from Behavioral Traces
Yi Qian, Fei Peng, Hao Wu, Ligeng Chen, Bing Mao
Abstract
We introduce prompt cloning, a new black-box attack that reconstructs functionally equivalent system prompts rather than extracts original system prompts. Unlike prompt stealing, prompt cloning exploits the insight that system prompts leave persistent behavioral traces in outputs, even under strong alignment and prompt-level defenses. Our method decomposes system behavior into semantically interpretable elements, selectively elicits them through carefully designed queries, and aggregates representative traces to synthesize high-fidelity cloned prompts. Extensive evaluations show that cloned prompts replicate functional behavior with up to 85% semantic similarity, outperforming base LLMs by up to 8%, and even exceeding original system prompts when transferred to different back-end models. We also conduct a large-scale study on GitHub repositories, revealing that single-prompt architectures remain widespread in open-source LLM applications, reinforcing the real-world relevance of our threat model. Our findings reveal that prompt cloning enables unauthorized replication of confidential LLM behavior and underscore the urgent need for defenses that go beyond hiding prompt text.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Prompt Obfuscation for Large Language ModelsDavid Pape, Sina Mavali, Thorsten Eisenhofer, Lea SchönherrUSENIX Security 2025
- PRSA: Prompt Stealing Attacks against Real-World Prompt ServicesYong Yang, Changjiang Li, Qingming Li, Oubo Ma et al.USENIX Security 2025
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM SystemsHongyan Chang, Ergute Bao, Xinjian Luo, Ting YuUSENIX Security 2026 · 24 citations
- Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language ModelsBiao Yi, Tiansheng Huang, Sishuo Chen, Tong Li et al.ICLR 2025
