GPTracker: A Large-Scale Measurement of Misused GPTs
Xinyue Shen, Yun Shen, Michael Backes, Yang Zhang
Abstract
Large language model (LLM)-powered agents, particularly GPTs by OpenAI, have revolutionized how AI is customized, deployed, and used. However, misuse of GPTs has emerged as a critical, yet largely underexplored, issue within OpenAI's GPT Store. In this paper, we present the first large-scale measurement study on misused GPTs. We introduce GPTRACKER, a framework designed to continuously collect GPTs from the official GPT Store and automate the interaction with them. As of the submission of this paper, GPTRACKER has collected 755,297 GPTs and 28,464 GPT conversation flows over eight months. Using an LLM-driven scoring system combined with human review, we identify 2,051 misused GPTs across ten forbidden scenarios. Through both static and dynamic analyses, we explore the landscape of these misused GPTs, including the trends, builders, operation mechanisms, and effectiveness. We find that builders of misused GPTs employ various tactics to bypass OpenAI's review system, such as integrating external APIs, hiding intention in descriptions, and URL redirection. Notably, GPTs activating external APIs are more likely to provide answers to inappropriate queries than other misused GPTs, showing an average 22.81% increase in answer rate in the Illegal Activity scenario. Leveraging VirusTotal, we identify 50 malicious domains shown on 446 GPTs, where 33 are labeled as phishing, 28 as malware, and 2 as spam, with some domains receiving multiple labels. We responsibly disclosed our findings to OpenAI on September 11, 2024, and November 12, 2024. 1,316 out of 1,804 GPTs reported in the first disclosure were removed by September 25. Our study sheds light on the alarming misuse of GPTs in the emerging GPT marketplace and offers actionable recommendations for stakeholders to mitigate future misuse.11Our code is available at https://github.com/TrustAIRLab/GPTracker. Disclaimer. This paper includes examples of hateful and disturbing content. Reader discretion is advised.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d463ecae-d1cb-4b7d-8d15-71042c496d78Cited by top-tier papers5
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the WildYi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng et al.USENIX Security 2026 · 46 citations
- Privacy in Human-AI Romantic Relationships: Concerns, Boundaries, and AgencyRongjun Ma, Shijing He, Jose Luis Martin-Navarro, Xiao Zhan et al.CHI 2026 · 5 citations
- When GPT Spills the Tea: Comprehensive Assessment of Knowledge File Leakage in GPTsXinyue Shen, Yun Shen, Michael Backes, Yang ZhangACL 2025
- ''Shut Up and Let Me Enjoy My Otome': Understanding and Measuring the Toxicity in Otome Game CommunitiesYage Zhang, Xinyue Shen, Yukun Jiang, Michael Backes et al.CCS 2026
- PROMPRINT: Prompt Fingerprinting via First-Token Response for LLM App Cloning DetectionJungmin Lee, Peizhuo Lv, Yeonjoon LeeACL 2026
Builds on15
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Formalizing and Benchmarking Prompt Injection Attacks and DefensesYupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia et al.USENIX Security 2024 · 308 citations
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen et al.NeurIPS 2024 · 195 citations
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen et al.CCS 2024 · 132 citations
Related papers
- On the (In)Security of LLM App StoresXinyi Hou, Yanjie Zhao, Haoyu WangS&P 2025
- From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language ModelsSayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, Shirin NilizadehS&P 2024 · 57 citations
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMsZhiyang Chen, Tara Saba, Xun Deng, Xujie Si et al.ICML 2026
- An Empirical Study to Evaluate AIGC Detectors on Code ContentJian Wang, Shangqing Liu, Xiaofei Xie, Yi LiASE 2024 · 4 citations
- Understanding and Detecting File Knowledge Leakage in GPT App EcosystemChuan Yan, Bowei Guan, Yazhi Li, Mark Huasong Meng et al.WWW 2025 · 5 citations
