PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning Detection
Jungmin Lee, Peizhuo Lv, Yeonjoon Lee
Abstract
As Large Language Model applications (LLM apps) become widespread, system prompts that determine app behavior are increasingly regarded as intellectual property, raising concerns about leakage. Recent studies show that this threat is no longer theoretical, revealing the prevalence of cloned apps replicating system prompts from others on real-world platforms. These clones pose risks of copyright infringement and malicious misuse, highlighting the need for early and reliable detection. In this paper, we propose PROMPRINT, a novel fingerprinting approach for detecting cloned LLM apps without exposing their system prompts. Motivated by the observation that different system prompts yield distinct first output token distributions for the same query, PROMPRINT optimizes queries that induce the LLM to generate a specific first output token associated with the given system prompt, resulting in distinctive query-first-token pairs. Experiments on four instruction-tuned LLMs show that generated pairs effectively identify the corresponding system prompts, achieving over 74% probability of generating the target token while remaining below 2.2% on average under other prompts. Furthermore, we demonstrate that our fingerprinting remains robust to partial system prompt modifications and effective under the injection of adversarial instructions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4d17713-1b7f-4819-a19f-df5875dbb665Builds on11
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
- FIRST: Faster Improved Listwise Reranking with Single Token DecodingRevanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md. Arafat Sultan et al.EMNLP 2024 · 14 citations
- The Curious Case of Hallucinatory (Un)answerability: Finding Truths in the Hidden States of Over-Confident Large Language ModelsAviv Slobodkin, Omer Goldman, Avi Caciularu, Ido Dagan et al.EMNLP 2023 · 11 citations
- Extracting Prompts by Inverting LLM OutputsCollin Zhang, John X. Morris, Vitaly ShmatikovEMNLP 2024 · 10 citations
- Understanding and Detecting File Knowledge Leakage in GPT App EcosystemChuan Yan, Bowei Guan, Yazhi Li, Mark Huasong Meng et al.WWW 2025 · 5 citations
Related papers
- Fingerprinting LLMs via Prompt InjectionYuepeng Hu, Zhengyuan Jiang, Mengyuan Li, Osama Ahmed et al.ACL 2026 · 3 citations
- PromptCOS: Towards Content-Only System Prompt Copyright Auditing for LLMsYuchen Yang, Yiming Li, Hongwei Yao, Enhao Huang et al.S&P 2026 · 5 citations
- Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting TechniqueMark Russinovich, Yanan Cai, Ahmed SalemICLR 2026 · 51 citations
- Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient EstimationShuo Shao, Yiming Li, Hongwei Yao, Yifei Chen et al.WWW 2026
- PromptCARE: Prompt Copyright Protection by Watermark Injection and VerificationHongwei Yao, Jian Lou, Zhan Qin, Kui RenS&P 2024 · 43 citations
