Fairshare Data Pricing via Data Valuation for Large Language Models
Luyang Zhang, Cathy Jiao, Beibei Li, Chenyan Xiong
Abstract
Training data is the backbone of large language models (LLMs), yet today's data markets often operate under exploitative pricing -- sourcing data from marginalized groups with little pay or recognition. This paper introduces a theoretical framework for LLM data markets, modeling the strategic interactions between buyers (LLM builders) and sellers (human annotators). We begin with theoretical and empirical analysis showing how exploitative pricing drives high-quality sellers out of the market, degrading data quality and long-term model performance. Then we introduce fairshare, a pricing mechanism grounded in data valuation that quantifies each data's contribution. It aligns incentives by sustaining seller participation and optimizing utility for both buyers and sellers. Theoretically, we show that fairshare yields mutually optimal outcomes: maximizing long-term buyer utility and seller profit while sustaining market participation. Empirically when training open-source LLMs on complex NLP tasks, including math problems, medical diagnosis, and physical reasoning, fairshare boosts seller earnings and ensures a stable supply of high-quality data, while improving buyers'performance-per-dollar and long-term welfare. Our findings offer a concrete path toward fair, transparent, and economically sustainable data markets for LLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b63d8fa-e99e-4ce3-b5a5-c60a75e47481Cited by top-tier papers3
- Modeling the Economic Impacts of AI Openness RegulationTori Qiu, Benjamin Laufer, Jon M. Kleinberg, Hoda HeidariNeurIPS 2025 · 5 citations
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
- Convex Dataset Valuation for Post-TrainingSiqi Zeng, Christopher Jung, Rui Li, Zhe Kang et al.ICML 2026
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
Related papers
- A Theory of Data Acquisition and Pricing at ScaleAndrew Ilyas, Amin Saberi, Grigorios VelegkasICML 2026
- Human-LLM Collaborative Annotation Through Effective Verification of LLM LabelsXinru Wang, Hannah Kim, Sajjadur Rahman, Kushan Mitra et al.CHI 2024 · 127 citations
- Pricing Online LLM Services with Data-Calibrated Stackelberg Routing GameZhendong Guo, Wenchao Bai, Jiahui JinAAAI 2026
- Incentivizing Quality Text Generation via Statistical ContractsEden Saig, Ohad Einav, Inbal Talgam-CohenNeurIPS 2024 · 18 citations
- Market-Bench: Benchmarking Large Language Models on Economic and Trade CompetitionYushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu et al.ACL 2026 · 1 citation
