Incentivizing Collaboration in Machine Learning via Synthetic Data Rewards
Sebastian Shenghong Tay, Xinyi Xu, Chuan Sheng Foo, Bryan Kian Hsiang Low
摘要
This paper presents a novel collaborative generative modeling (CGM) framework that incentivizes collaboration among self-interested parties to contribute data to a pool for training a generative model (e.g., GAN), from which synthetic data are drawn and distributed to the parties as rewards commensurate to their contributions. Distributing synthetic data as rewards (instead of trained models or money) offers task- and model-agnostic benefits for downstream learning tasks and is less likely to violate data privacy regulation. To realize the framework, we firstly propose a data valuation function using maximum mean discrepancy (MMD) that values data based on its quantity and quality in terms of its closeness to the true data distribution and provide theoretical results guiding the kernel choice in our MMD-based data valuation function. Then, we formulate the reward scheme as a linear optimization problem that when solved, guarantees certain incentives such as fairness in the CGM framework. We devise a weighted sampling algorithm for generating synthetic data to be distributed to each party as reward such that the value of its data and the synthetic data combined matches its assigned reward value by the reward scheme. We empirically show using simulated and real-world datasets that the parties' synthetic data rewards are commensurate to their contributions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Gradient Driven Rewards to Guarantee Fairness in Collaborative Machine LearningXinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao 等NeurIPS 2021 · 被引用 133 次
- Validation Free and Replication Robust Volume-based Data ValuationXinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang LowNeurIPS 2021 · 被引用 89 次
- DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang LowICML 2022 · 被引用 71 次
- Secure Shapley Value for Cross-Silo Federated LearningShuyuan Zheng, Yang Cao, Masatoshi YoshikawaVLDB 2023 · 被引用 41 次
- Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon 等ICML 2024 · 被引用 24 次
它引用的顶会 Paper6
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 被引用 158 次
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 被引用 152 次
- If You Like Shapley Then You'll Love the CoreTom Yan, Ariel D. ProcacciaAAAI 2021 · 被引用 85 次
- A Scalable Approach for Privacy-Preserving Collaborative Machine LearningJinhyun So, Basak Güler, Salman AvestimehrNeurIPS 2020 · 被引用 60 次
- Differentially Private and Communication Efficient Collaborative LearningJiahao Ding, Guannan Liang, Jinbo Bi, Miao PanAAAI 2021 · 被引用 28 次
相关 Paper
- Collaborative Causal Inference with Fair IncentivesRui Qiao, Xinyi Xu, Bryan Kian Hsiang LowICML 2023 · 被引用 8 次
- Incentives in Private Collaborative Machine LearningRachael Hwee Ling Sim, Yehong Zhang, Nghia Hoang, Xinyi Xu 等NeurIPS 2023 · 被引用 12 次
- GMValuator: Similarity-based Data Valuation for Generative ModelsJiaxi Yang, Wenlong Deng, Benlin Liu, Yangsibo Huang 等ICLR 2025
- A Cramér-von Mises Approach to Incentivizing Truthful Data SharingAlex Clinton, Thomas Zeng, Yiding Chen, Xiaojin Zhu 等NeurIPS 2025 · 被引用 2 次
- A Characteristic Function Approach to Deep Implicit Generative ModelingAbdul Fatir Ansari, Jonathan Scarlett, Harold SohCVPR 2020
