Lune

ASE2025顶会

Uncovering Prompt Elements: Cloning System Prompts from Behavioral Traces

Yi Qian, Fei Peng, Hao Wu, Ligeng Chen, Bing Mao

2025年份
1被引次数

摘要

We introduce prompt cloning, a new black-box attack that reconstructs functionally equivalent system prompts rather than extracts original system prompts. Unlike prompt stealing, prompt cloning exploits the insight that system prompts leave persistent behavioral traces in outputs, even under strong alignment and prompt-level defenses. Our method decomposes system behavior into semantically interpretable elements, selectively elicits them through carefully designed queries, and aggregates representative traces to synthesize high-fidelity cloned prompts. Extensive evaluations show that cloned prompts replicate functional behavior with up to 85% semantic similarity, outperforming base LLMs by up to 8%, and even exceeding original system prompts when transferred to different back-end models. We also conduct a large-scale study on GitHub repositories, revealing that single-prompt architectures remain widespread in open-source LLM applications, reinforcing the real-world relevance of our threat model. Our findings reveal that prompt cloning enables unauthorized replication of confidential LLM behavior and underscore the urgent need for defenses that go beyond hiding prompt text.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖