A Cramér-von Mises Approach to Incentivizing Truthful Data Sharing
Alex Clinton, Thomas Zeng, Yiding Chen, Xiaojin Zhu, Kirthevasan Kandasamy
Abstract
Modern data marketplaces and data sharing consortia increasingly rely on incentive mechanisms to encourage agents to contribute data. However, schemes that reward agents based on the quantity of submitted data are vulnerable to manipulation, as agents may submit fabricated or low-quality data to inflate their rewards. Prior work has proposed comparing each agent's data against others' to promote honesty: when others contribute genuine data, the best way to minimize discrepancy is to do the same. Yet prior implementations of this idea rely on very strong assumptions about the data distribution (e.g. Gaussian), limiting their applicability. In this work, we develop reward mechanisms based on a novel two-sample test statistic inspired by the Cramér-von Mises statistic. Our methods strictly incentivize agents to submit more genuine data, while disincentivizing data fabrication and other types of untruthful reporting. We establish that truthful reporting constitutes a (possibly approximate) Nash equilibrium in both Bayesian and prior-agnostic settings. We theoretically instantiate our method in three canonical data sharing problems and show that it relaxes key assumptions made by prior work. Empirically, we demonstrate that our mechanism incentivizes truthful data sharing via simulations and on real-world language and image data. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d9dc49a-66ad-407d-9cd6-65fc5ec7a991Builds on8
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- One for One, or All for All: Equilibria and Optimality of Collaboration in Federated LearningAvrim Blum, Nika Haghtalab, Richard Lanas Phillips, Han ShaoICML 2021 · 62 citations
- Truthful Data Acquisition via Peer PredictionYiling Chen, Yiheng Shen, Shuran ZhengNeurIPS 2020 · 35 citations
- Incentivizing Honesty among Competitors in Collaborative Learning and OptimizationFlorian E. Dorner, Nikola Konstantinov, Georgi Pashaliev, Martin T. VechevNeurIPS 2023 · 18 citations
Related papers
- Incentivizing Collaboration in Machine Learning via Synthetic Data RewardsSebastian Shenghong Tay, Xinyi Xu, Chuan Sheng Foo, Bryan Kian Hsiang LowAAAI 2022 · 40 citations
- Incentivizing Truthfulness and Collaborative Fairness in Bayesian LearningRachael Hwee Ling Sim, Jue Fan, Xiao Tian, Xinyi Xu et al.ICML 2026
- Mechanism Design for Collaborative Normal Mean EstimationYiding Chen, Jerry Zhu, Kirthevasan KandasamyNeurIPS 2023 · 15 citations
- Collaborative Causal Inference with Fair IncentivesRui Qiao, Xinyi Xu, Bryan Kian Hsiang LowICML 2023 · 8 citations
- Preventing Strategic Behaviors in Collaborative Inference for Vertical Federated LearningYidan Xing, Zhenzhe Zheng, Fan WuKDD 2024 · 1 citation
