Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling Algorithm
Boxin Zhao, Boxiang Lyu, Raul Castro Fernandez, Mladen Kolar
摘要
High-quality machine learning models are dependent on access to high-quality training data. When the data are not already available, it is tedious and costly to obtain them. Data markets help with identifying valuable training data: model consumers pay to train a model, the market uses that budget to identify data and train the model (the budget allocation problem), and finally the market compensates data providers according to their data contribution (revenue allocation problem). For example, a bank could pay the data market to access data from other financial institutions to train a fraud detection model. Compensating data contributors requires understanding data's contribution to the model; recent efforts to solve this revenue allocation problem based on the Shapley value are inefficient to lead to practical data markets. In this paper, we introduce a new algorithm to solve budget allocation and revenue allocation problems simultaneously in linear time. The new algorithm employs an adaptive sampling process that selects data from those providers who are contributing the most to the model. Better data means that the algorithm accesses those providers more often, and more frequent accesses corresponds to higher compensation. Furthermore, the algorithm can be deployed in both centralized and federated scenarios, boosting its applicability. We provide theoretical guarantees for the algorithm that show the budget is used efficiently and the properties of revenue allocation are similar to Shapley's. Finally, we conduct an empirical evaluation to show the performance of the algorithm in practical scenarios and when compared to other baselines. Overall, we believe that the new algorithm paves the way for the implementation of practical data markets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao 等NeurIPS 2025 · 被引用 112 次
- Data Acquisition via Experimental Design for Data MarketsCharles Lu, Baihe Huang, Sai Praneeth Karimireddy, Praneeth Vepakomma 等NeurIPS 2024 · 被引用 12 次
- 2D-OOB: Attributing Data Contribution Through Joint Valuation FrameworkYifan Sun, Jingyan Shen, Yongchan KwonNeurIPS 2024 · 被引用 8 次
- Data Acquisition for Improving Model ConfidenceYifan Li, Xiaohui Yu, Nick KoudasSIGMOD 2024 · 被引用 4 次
- How to Price Data: A Market Equilibrium Based ApproachPooja Kulkarni, Parnian Shahkar, Ruta MehtaICML 2026
它引用的顶会 Paper9
- SCAFFOLD: Stochastic Controlled Averaging for Federated LearningSai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J. Reddi 等ICML 2020 · 被引用 3,875 次
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 被引用 236 次
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 被引用 152 次
- Dealer: An End-to-End Model Marketplace with Differential PrivacyJinfei Liu, Jian Lou, Junxu Liu, Li Xiong 等VLDB 2021 · 被引用 99 次
- Revenue Maximization for Query PricingShuchi Chawla, Shaleen Deep, Paraschos Koutris, Yifeng TengVLDB 2020 · 被引用 58 次
相关 Paper
- On Shapley Value in Data Assemblage Under Independent UtilityXuan Luo, Jian Pei, Zicun Cong, Cheng XuVLDB 2022 · 被引用 18 次
- Data-Sharing Markets: Model, Protocol, and Algorithms to Incentivize the Formation of Data-Sharing ConsortiaRaul Castro FernandezSIGMOD 2023 · 被引用 28 次
- Equitable Data Valuation Meets the Right to Be Forgotten in Model MarketsHaocheng Xia, Jinfei Liu, Jian Lou, Zhan Qin 等VLDB 2023 · 被引用 28 次
- Faithful Group Shapley ValueKiljae Lee, Ziqi Liu, Weijing Tang, Yuan ZhangNeurIPS 2025 · 被引用 4 次
- EcoVal: An Efficient Data Valuation Framework for Machine LearningAyush K. Tarun, Vikram S. Chundawat, Murari Mandal, Hong Ming Tan 等KDD 2024 · 被引用 3 次
