Real Money, Fake Models: Deceptive Model Claims in Shadow APIs
Yage Zhang, Yukun Jiang, Zeyuan Chen, Michael Backes, Xinyue Shen, Yang Zhang
摘要
Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions. These limitations drive the proliferation of , third-party services that claim to provide access to official model services without regional limitations via indirect access. Despite their widespread use, it remains unclear whether shadow APIs deliver outputs consistent with those of the official APIs, raising concerns about the reliability of downstream applications and the validity of research findings that depend on them. In this paper, we present the first systematic audit between official LLM APIs and corresponding shadow APIs. We first identify 17 shadow APIs that have been utilized in 187 academic papers, with the most popular one reaching more than 5,900 citations and 58,000 GitHub stars by December 6, 2025. Through multidimensional auditing of three representative shadow APIs across utility, safety, and model verification, we uncover widespread behavioral inconsistency and fingerprint-based evidence consistent with deceptive model claims in a subset of audited endpoints. Specifically, we reveal performance divergence reaching up to 47.21%, significant unpredictability in safety behaviors, and identity verification failures in 45.83% of fingerprint tests. These practices critically undermine the reproducibility and validity of scientific research, harm the interests of shadow API users, and damage the reputation of official model providers. By the time of writing, 4 of the 17 providers have already ceased operations, underscoring the operational volatility of this market. Meanwhile, unverifiable compliance claims and independent model-substitution testing platforms have emerged in the ecosystem, reflecting growing community awareness of this risk.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability CurvatureGuangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang 等ICLR 2024 · 被引用 311 次
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot 等ICLR 2020 · 被引用 244 次
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold 等ICLR 2024 · 被引用 173 次
相关 Paper
- Black-Box Detection of Language Model WatermarksThibaud Gloaguen, Nikola Jovanovic, Robin Staab, Martin T. VechevICLR 2025
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- PaperAsk: A Benchmark for Reliability Evaluation of LLMs in Paper Search and ReadingYutao Wu, Xiao Liu, Yunhao Feng, Jiale Ding 等WWW 2026 · 被引用 1 次
- Malla: Demystifying Real-world Large Language Model Integrated Malicious ServicesZilong Lin, Jian Cui, Xiaojing Liao, XiaoFeng WangUSENIX Security 2024 · 被引用 49 次
- On the Reliability of Psychological Scales on Large Language ModelsJen-tse Huang, Wenxiang Jiao, Man Ho Lam, Eric John Li 等EMNLP 2024 · 被引用 6 次
