Validation Free and Replication Robust Volume-based Data Valuation
Xinyi Xu, Zhaoxuan Wu, Chuan Sheng Foo, Bryan Kian Hsiang Low
摘要
Data valuation arises as a non-trivial challenge in real-world use cases such as collaborative machine learning, federated learning, trusted data sharing, data marketplaces. The value of data is often associated with the learning performance (e.g., validation accuracy) of a model trained on the data, which introduces a close coupling between data valuation and validation. However, a validation set may not be available in practice and it can be challenging for the data providers to reach an agreement on the choice of the validation set. Another practical issue is that of data replication: Given the value of some data points, a dishonest data provider may replicate these data points to exploit the valuation for a larger reward/payment. We observe that the diversity of the data points is an inherent property of a dataset that is independent of validation. We formalize diversity via the volume of the data matrix (i.e., determinant of its left Gram), which allows us to establish a formal connection between the diversity of data and learning performance without requiring validation. Furthermore, we propose a robust volume measure with a theoretical guarantee on the replication robustness by following the intuition that copying the same data points does not increase the diversity of data. We perform extensive experiments to demonstrate its consistency in valuation and practical advantages over existing baselines and show that our method is model-and task-agnostic and can be flexibly adapted to handle various neural networks. * Equal contribution. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu 等ICLR 2024 · 被引用 281 次
- Gradient Driven Rewards to Guarantee Fairness in Collaborative Machine LearningXinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao 等NeurIPS 2021 · 被引用 133 次
- Fault-Tolerant Federated Reinforcement Learning with Theoretical GuaranteeFlint Xiaofeng Fan, Yining Ma, Zhongxiang Dai, Wei Jing 等NeurIPS 2021 · 被引用 102 次
- DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang LowICML 2022 · 被引用 71 次
- Rethinking Data Shapley for Data Selection Tasks: Misleads and MeritsJiachen T. Wang, Tianji Yang, James Zou, Yongchan Kwon 等ICML 2024 · 被引用 24 次
它引用的顶会 Paper5
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 被引用 236 次
- Collaborative Machine Learning with Incentive-Aware Model RewardsRachael Hwee Ling Sim, Yehong Zhang, Mun Choon Chan, Bryan Kian Hsiang LowICML 2020 · 被引用 158 次
- Gradient Driven Rewards to Guarantee Fairness in Collaborative Machine LearningXinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao 等NeurIPS 2021 · 被引用 133 次
- Incentivizing Collaboration in Machine Learning via Synthetic Data RewardsSebastian Shenghong Tay, Xinyi Xu, Chuan Sheng Foo, Bryan Kian Hsiang LowAAAI 2022 · 被引用 40 次
- Influence Functions in Deep Learning Are FragileSamyadeep Basu, Phillip Pope, Soheil FeiziICLR 2021 · 被引用 15 次
相关 Paper
- Distributionally Robust Data ValuationXiaoqiang Lin, Xinyi Xu, Zhaoxuan Wu, See-Kiong Ng 等ICML 2024 · 被引用 6 次
- Fundamentals of Task-Agnostic Data ValuationMohammad Mohammadi Amiri, Frederic Berdoz, Ramesh RaskarAAAI 2023 · 被引用 20 次
- LAVA: Data Valuation without Pre-Specified Learning AlgorithmsHoang Anh Just, Feiyang Kang, Tianhao Wang, Yi Zeng 等ICLR 2023 · 被引用 6 次
- Improving Fairness for Data Valuation in Horizontal Federated LearningZhenan Fan, Huang Fang, Zirui Zhou, Jian Pei 等ICDE 2022 · 被引用 68 次
- Fortifying Federated Learning Towards Trustworthiness via Auditable Data Valuation and Verifiable Client ContributionK. Naveen Kumar, Ranjeet Ranjan Jha, C. Krishna Mohan, Ravindra Babu TallamrajuCVPR 2025
