Incentivizing Collaboration in Machine Learning via Synthetic Data Rewards
Sebastian Shenghong Tay, Xinyi Xu, Chuan Sheng Foo, Bryan Kian Hsiang Low
Abstract
This paper presents a novel collaborative generative modeling (CGM) framework that incentivizes collaboration among self-interested parties to contribute data to a pool for training a generative model (e.g., GAN), from which synthetic data are drawn and distributed to the parties as rewards commensurate to their contributions. Distributing synthetic data as rewards (instead of trained models or money) offers task- and model-agnostic benefits for downstream learning tasks and is less likely to violate data privacy regulation. To realize the framework, we firstly propose a data valuation function using maximum mean discrepancy (MMD) that values data based on its quantity and quality in terms of its closeness to the true data distribution and provide theoretical results guiding the kernel choice in our MMD-based data valuation function. Then, we formulate the reward scheme as a linear optimization problem that when solved, guarantees certain incentives such as fairness in the CGM framework. We devise a weighted sampling algorithm for generating synthetic data to be distributed to each party as reward such that the value of its data and the synthetic data combined matches its assigned reward value by the reward scheme. We empirically show using simulated and real-world datasets that the parties' synthetic data rewards are commensurate to their contributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2fa4a40-6cd2-4809-ac94-5282aa9556bfCited by top-tier papers23
- Gradient Driven Rewards to Guarantee Fairness in Collaborative Machine LearningXinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao et al.NeurIPS 2021 · 133 citations
- Validation Free and Replication Robust Volume-based Data ValuationXinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang LowNeurIPS 2021 · 89 citations
- DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang LowICML 2022 · 71 citations
- Secure Shapley Value for Cross-Silo Federated LearningShuyuan Zheng, Yang Cao, Masatoshi YoshikawaVLDB 2023 · 41 citations
- Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon et al.ICML 2024 · 24 citations
Builds on6
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 158 citations
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 152 citations
- If You Like Shapley Then You'll Love the CoreTom Yan, Ariel D. ProcacciaAAAI 2021 · 85 citations
- A Scalable Approach for Privacy-Preserving Collaborative Machine LearningJinhyun So, Basak Güler, Salman AvestimehrNeurIPS 2020 · 60 citations
- Differentially Private and Communication Efficient Collaborative LearningJiahao Ding, Guannan Liang, Jinbo Bi, Miao PanAAAI 2021 · 28 citations
Related papers
- Collaborative Causal Inference with Fair IncentivesRui Qiao, Xinyi Xu, Bryan Kian Hsiang LowICML 2023 · 8 citations
- Incentives in Private Collaborative Machine LearningRachael Hwee Ling Sim, Yehong Zhang, Nghia Hoang, Xinyi Xu et al.NeurIPS 2023 · 12 citations
- GMValuator: Similarity-based Data Valuation for Generative ModelsJiaxi Yang, Wenlong Deng, Benlin Liu, Yangsibo Huang et al.ICLR 2025
- A Cramér-von Mises Approach to Incentivizing Truthful Data SharingAlex Clinton, Thomas Zeng, Yiding Chen, Xiaojin Zhu et al.NeurIPS 2025 · 2 citations
- A Characteristic Function Approach to Deep Implicit Generative ModelingAbdul Fatir Ansari, Jonathan Scarlett, Harold SohCVPR 2020
