Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
Ander Artola Velasco, Stratis Tsirtsis, Nastaran Okati, Manuel Gomez-Rodriguez
摘要
State-of-the-art large language models require specialized hardware and substantial energy to operate. Consequently, cloud-based services that provide access to these models have become very popular. In these services, the price users pay depends on the number of tokens a model uses to generate an output–they pay a fixed price per token. In this work, we show that this pricing mechanism creates a financial incentive for providers to strategize and misreport the (number of) tokens a model used to generate an output, and users cannot prove, or even know, whether a provider is overcharging them. However, we also show that, if an unfaithful provider is obliged to be transparent about the generative process used by the model, misreporting optimally without raising suspicion is hard. Nevertheless, as a proof-of-concept, we develop an efficient heuristic algorithm that allows providers to significantly overcharge users without raising suspicion. Crucially, the cost of running the algorithm is lower than the additional revenue from overcharging users, highlighting the vulnerability of users under the current pay-per-token pricing mechanism. Further, we show that, to eliminate the financial incentive to strategize, a pricing mechanism must price tokens linearly on their character count. While this makes a provider's profit margin vary across tokens, we introduce a simple prescription that allows a provider to maintain their average profit margin when transitioning to an incentive-compatible pricing mechanism. To complement our theoretical results, we conduct experiments with large language models from the , and families, and prompts from a popular benchmarking platform.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Modeling the Economic Impacts of AI Openness RegulationTori Qiu, Benjamin Laufer, Jon M. Kleinberg, Hoda HeidariNeurIPS 2025 · 被引用 5 次
- Pay for The Second-Best Service: A Game-Theoretic Approach against Dishonest LLM ProvidersYuhan Cao, Yu Wang, Sitong Liu, Miao Li 等WWW 2026 · 被引用 3 次
- Adaptive Contracts for Cost-Effective AI DelegationEden Saig, Tamar Garbuz, Ariel Procaccia, Inbal Talgam-Cohen 等ICML 2026 · 被引用 2 次
- Computational Arbitrage in AI Model MarketsRicardo Dominguez-Olmedo, Bernhard Schölkopf, Moritz HardtICML 2026
它引用的顶会 Paper20
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
- Language Model Tokenizers Introduce Unfairness Between LanguagesAleksandar Petrov, Emanuele La Malfa, Philip H. S. Torr, Adel BibiNeurIPS 2023 · 被引用 301 次
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
- Mechanism Design for Large Language ModelsPaul Dütting, Vahab Mirrokni, Renato Paes Leme, Haifeng Xu 等WWW 2024 · 被引用 65 次
相关 Paper
- TokenPowerBench: Benchmarking the Power Consumption of LLM InferenceChenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao 等AAAI 2026 · 被引用 12 次
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-Based Prompt Injection Attacks via the Fine-Tuning InterfaceAndrey Labunets, Nishit V. Pandya, Ashish Hooda, Xiaohan Fu 等S&P 2025
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 被引用 17 次
- The Invisible Hand: Unveiling Provider Bias in Large Language Models for Code GenerationXiaoyu Zhang, Juan Zhai, Shiqing Ma, Qingshuang Bao 等ACL 2025 · 被引用 6 次
- An Engorgio Prompt Makes Large Language Model Babble onJianshuo Dong, Ziyuan Zhang, Qingjie Zhang, Tianwei Zhang 等ICLR 2025
