Data-Sharing Markets: Model, Protocol, and Algorithms to Incentivize the Formation of Data-Sharing Consortia
Raul Castro Fernandez
摘要
Organizations that would mutually benefit from pooling their data are otherwise wary of sharing. This is because sharing data is costly-in time and effort-and, at the same time, the benefits of sharing are not clear. Without a clear cost-benefit analysis, participants default in not sharing. As a consequence, many opportunities to create valuable data-sharing consortia never materialize and the value of data remains locked. We introduce a new sharing model, market protocol, and algorithms to incentivize the creation of data-sharing markets. The combined contributions of this paper, which we call DSC, incentivize the creation of data-sharing markets that unleash the value of data for its participants. The sharing model introduces two incentives; one that guarantees that participating is better than not doing so, and another that compensates participants according to how valuable is their data. Because operating the consortia is costly, we are also concerned with ensuring its operation is sustainable: we design a protocol that ensures that valuable data-sharing consortia form when it is sustainable. We introduce algorithms to elicit the value of data from the participants, which is used to: first, cover the costs of operating the consortia, and second compensate data contributions. For the latter, we challenge the use of the Shapley value to allocate revenue. We offer analytical and empirical evidence for this and introduce an alternative method that compensates participants better and leads to the formation of more data-sharing consortia.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsSang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao 等NeurIPS 2025 · 被引用 112 次
- Data Acquisition via Experimental Design for Data MarketsCharles Lu, Baihe Huang, Sai Praneeth Karimireddy, Praneeth Vepakomma 等NeurIPS 2024 · 被引用 12 次
- Fast, Robust and Interpretable Participant Contribution Estimation for Federated LearningYong Wang, Kaiyu Li, Yuyu Luo, Guoliang Li 等ICDE 2024 · 被引用 10 次
- Evaluating the Privacy Valuation of Personal Data on SmartphonesLihua Fan, Shuning Zhang, Yan Kong, Xin Yi 等UbiComp 2024 · 被引用 4 次
- Data Acquisition for Improving Model ConfidenceYifan Li, Xiaohui Yu, Nick KoudasSIGMOD 2024 · 被引用 4 次
它引用的顶会 Paper8
- Machine Learning Models that Remember Too MuchCongzheng Song, Thomas Ristenpart, Vitaly ShmatikovCCS 2017 · 被引用 582 次
- Lessons Learned from the Chameleon TestbedKate Keahey, Jason Anderson, Zhuo Zhen, Pierre Riteau 等USENIX ATC 2020 · 被引用 398 次
- An Incentive Mechanism for Cross-Silo Federated Learning: A Public Goods PerspectiveMing Tang, Vincent W. S. WongINFOCOM 2021 · 被引用 122 次
- Hu-Fu: Efficient and Secure Spatial Queries over Data FederationYongxin Tong, Xuchen Pan, Yuxiang Zeng, Yexuan Shi 等VLDB 2022 · 被引用 63 次
- Data Acquisition for Improving Machine Learning ModelsYifan Li, Xiaohui Yu, Nick KoudasVLDB 2021 · 被引用 57 次
相关 Paper
- Addressing Budget Allocation and Revenue Allocation in Data Market Environments Using an Adaptive Sampling AlgorithmBoxin Zhao, Boxiang Lyu, Raul Castro Fernandez, Mladen KolarICML 2023 · 被引用 14 次
- On Shapley Value in Data Assemblage Under Independent UtilityXuan Luo, Jian Pei, Zicun Cong, Cheng XuVLDB 2022 · 被引用 18 次
- Collaborative Causal Inference with Fair IncentivesRui Qiao, Xinyi Xu, Bryan Kian Hsiang LowICML 2023 · 被引用 8 次
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 被引用 152 次
- Equitable Data Valuation Meets the Right to Be Forgotten in Model MarketsHaocheng Xia, Jinfei Liu, Jian Lou, Zhan Qin 等VLDB 2023 · 被引用 28 次
