LLMThief: Evaluating Configuration Leaking Risks in Commercial LLM App Stores
Pinji Chen, Jinlong Jiang, Jianjun Chen, Feiran Qin, Minghao Zhang, Jiahe Zhang, Haixin Duan, Kaiwen Shen, Hui Jiang
Abstract
Configuration leaking attack is an emerging security threat in large language model applications (LLM apps), where adversaries can manipulate the LLM app to reveal its sensitive configurations, such as system prompts, external APIs, and knowledge files. Despite their critical implications, these attacks remain understudied within commercial LLM app stores, leaving open questions about their real-world effectiveness, prevalence, and impacts. In this paper, we propose LLMThief, a novel end-to-end framework designed to systematically evaluate configuration leaking risks in LLM app stores. LLMThief comprises three phases. First, it extracts grounding information from store-level features to construct a high-quality seed pool of attack prompts. Second, it leverages shadow LLM apps as probing oracles and a genetic algorithm to identify adaptive mutation strategies. Third, it fuzzes victim LLM apps with the resulting insightful prompts and uses a fine-tuned LLM to determine whether a configuration leakage has occurred.
We evaluated LLMThief on ground truth datasets across 6 widely used LLM app stores, including OpenAI GPT Store, ByteDance Coze, and Baidu Wenxin. Evaluation results show that LLMThief can effectively leak the confidential configuration of LLM apps and significantly outperforms baselines. Beyond performance evaluation, our large-scale analysis of 4,164 real-world LLM apps reveals a range of critical risks, including system prompt leaks, external API exposures, and knowledge file leaks. Those issues can not only compromise developers' intellectual property but also leak personal privacy and even disclose corporate secrets. We have responsibly disclosed our findings to the affected vendors and received acknowledgments and bug bounties from Baidu, ByteDance, Alibaba, etc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- How Language Model Hallucinations Can SnowballMuru Zhang, Ofir Press, William Merrill, Alisa Liu et al.ICML 2024 · 406 citations
- One-Shot Safety Alignment for Large Language Models via Optimal DualizationXinmeng Huang, Shuo Li, Edgar Dobriban, Osbert Bastani et al.NeurIPS 2024 · 30 citations
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
Related papers
- PRSA: Prompt Stealing Attacks against Real-World Prompt ServicesYong Yang, Changjiang Li, Qingming Li, Oubo Ma et al.USENIX Security 2025
- Mimicking the Familiar: Dynamic Command Generation for Information Theft Attacks in LLM Tool-Learning SystemZiyou Jiang, Mingyang Li, Guowei Yang, Junjie Wang et al.ACL 2025
- Understanding and Mitigating Prompt Leaking Attacks in Real-World LLM-Based ApplicationsYong Yang, Chong Fu, Tong Zhang, Rui Zeng et al.CCS 2026
- RepeatLeakage: Leak Prompts from Repeating as Large Language Model Is a Good RepeaterYu Peng, Lijie Zhang, Peizhuo Lv, Kai ChenAAAI 2025 · 3 citations
- TrojLLM: A Black-box Trojan Prompt Attack on Large Language ModelsJiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen et al.NeurIPS 2023 · 63 citations
