On the (In)Security of LLM App Stores
Xinyi Hou, Yanjie Zhao, Haoyu Wang
摘要
LLM app stores have seen rapid growth, leading to the proliferation of numerous custom LLM apps. However, this expansion raises security concerns. In this study, we propose a three-layer concern framework to identify the potential security risks of LLM apps, i.e., LLM apps with abusive potential, LLM apps with malicious intent, and LLM apps with backdoors. Over five months, we collected 786,036 LLM apps from six major app stores: GPT Store, FlowGPT, Poe, Coze, Cici, and Character.AI. Our research integrates static and dynamic analysis, and uses a complementary approach to detect harmful content, combining a self-refining LLM-based toxic content detector with rule-based pattern matching. Additionally, we constructed a large-scale toxic word dictionary (i.e., ToxicDict) comprising over 31,783 entries. We used these methods to uncover that 15,414 apps had misleading descriptions, 1,366 collected sensitive personal information against their privacy policies, and 15,996 generated harmful content such as hate speech, self-harm, extremism, etc. Additionally, we evaluated the potential for LLM apps to facilitate malicious activities, finding that 616 apps could be used for malware generation, phishing, etc. We reported these security risks to relevant platforms, including OpenAI and Quora, which acknowledged and appreciated our findings. The platforms are actively investigating the flagged apps; as of the submission of this paper, 1,643 apps have been removed from the GPT Store.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Privacy Perceptions of Custom GPTs by Users and CreatorsRongjun Ma, Caterina Maidhof, Juan Carlos Carrillo, Janne Lindqvist 等CHI 2025 · 被引用 35 次
- PromptCOS: Towards Content-Only System Prompt Copyright Auditing for LLMsYuchen Yang, Yiming Li, Hongwei Yao, Enhao Huang 等S&P 2026 · 被引用 5 次
- Understanding and Detecting File Knowledge Leakage in GPT App EcosystemChuan Yan, Bowei Guan, Yazhi Li, Mark Huasong Meng 等WWW 2025 · 被引用 5 次
- LLMThief: Evaluating Configuration Leaking Risks in Commercial LLM App StoresPinji Chen, Jinlong Jiang, Jianjun Chen, Feiran Qin 等S&P 2026 · 被引用 1 次
- Malicious LLM-Based Conversational AI Makes Users Reveal Personal InformationXiao Zhan, Juan Carlos Carrillo, William Seymour, Jose SuchUSENIX Security 2025
它引用的顶会 Paper5
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Malla: Demystifying Real-world Large Language Model Integrated Malicious ServicesZilong Lin, Jian Cui, Xiaojing Liao, XiaoFeng WangUSENIX Security 2024 · 被引用 49 次
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina 等CCS 2024 · 被引用 28 次
相关 Paper
- GPTracker: A Large-Scale Measurement of Misused GPTsXinyue Shen, Yun Shen, Michael Backes, Yang ZhangS&P 2025
- Efficient Detection of Toxic Prompts in Large Language ModelsYi Liu, Junzhe Yu, Huijia Sun, Ling Shi 等ASE 2024 · 被引用 6 次
- Exploring ChatGPT App Ecosystem: Distribution, Deployment and SecurityChuan Yan, Ruomai Ren, Mark Huasong Meng, Liuhuo Wan 等ASE 2024 · 被引用 3 次
- On Large Language Models' Resilience to Coercive InterrogationZhuo Zhang, Guangyu Shen, Guanhong Tao, Siyuan Cheng 等S&P 2024 · 被引用 24 次
- Coverage-Based Harmfulness Testing for LLM Code TransformationHonghao Tan, Haibo Wang, Diany Pressato, Yisen Xu 等ASE 2025 · 被引用 2 次
