Cost-of-Pass: An Economic Framework for Evaluating Language Models
Mehmet Hamza Erol, Batu El, Mirac Suzgun, Mert Yüksekgönül, James Y. Zou
Abstract
The widespread adoption of AI systems in the economy hinges on their ability to generate economic value that outweighs their inference costs. Evaluating this tradeoff requires metrics that account for both performance and costs. Building on Farrell's theory of productive efficiency, we develop an economically grounded framework for evaluating language models' productivity by combining accuracy and inference cost. We formalize cost-of-pass, the expected monetary cost of generating a correct solution. We then define the frontier cost-of-pass as the minimum cost-of-pass achievable across available models or the human-expert(s), using the approximate cost of hiring an expert. Our analysis reveals distinct economic insights. First, lightweight models are most cost-effective for basic quantitative tasks, large models for knowledge-intensive ones, and reasoning models for complex quantitative problems, despite higher per-token costs. Second, tracking this frontier cost-of-pass over the past year reveals significant progress, particularly for complex quantitative tasks where the cost has roughly halved every few months. Third, to trace key innovations driving this progress, we examine counterfactual frontiers-estimates of cost-efficiency without specific model classes. We find that innovations in lightweight, large, and reasoning models have been essential for pushing the frontier in basic quantitative, knowledge-intensive, and complex quantitative tasks, respectively. Finally, we assess the cost-reductions from common inference-time techniques (majority voting and self-refinement), and a budgetaware technique (TALE-EP). We find that performance-oriented methods with marginal performance gains rarely justify the costs, while TALE-EP shows some promise. Overall, our findings underscore that complementary model-level innovations are the primary drivers of cost-efficiency, and our economic framework provides a principled tool for measuring this progress and guiding deployment. *
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- AFM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid ReasoningQianben Chen, Jingyi Cao, Jiayu Zhang, Tianrui Qin et al.ICLR 2026 · 3 citations
- Reasoning Language Model Inference Serving Unveiled: An Empirical StudyQi Li, Junpan Wu, Xiang Liu, Yuxin Wang et al.ICLR 2026 · 3 citations
- Routing, Cascades, and User Choice for LLMsRafid MahmoodICLR 2026 · 2 citations
- Computational Arbitrage in AI Model MarketsRicardo Dominguez-Olmedo, Bernhard Schölkopf, Moritz HardtICML 2026
Builds on10
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Training Language Models to Reason EfficientlyDaman Arora, Andrea ZanetteNeurIPS 2025 · 270 citations
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
Related papers
- EconProver: Towards More Economical Test-Time Scaling for Automated Theorem ProvingMukai Li, Linfeng Song, Zhenwen Liang, Jiahao Xu et al.ACL 2026
- Plan and Budget: Effective and Efficient Test-Time Scaling on Reasoning Large Language ModelsJunhong Lin, Xinyue Zeng, Jie Zhu, Song Wang et al.ICLR 2026 · 30 citations
- Cheaply Estimating Inference Efficiency Metrics for Autoregressive Transformer ModelsDeepak Narayanan, Keshav Santhanam, Peter Henderson, Rishi Bommasani et al.NeurIPS 2023 · 14 citations
- Tina: Tiny Reasoning Models via LoRAShangshang Wang, Julian Asilis, Ömer Faruk Akgül, Enes Burak Bilgin et al.ICLR 2026 · 30 citations
- Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Significant Gains in Reasoning Efficiency in Large Language ModelsQiguang Chen, Dengyun Peng, Jinhao Liu, Huikang Su et al.AAAI 2026
