How much of my dataset did you use? Quantitative Data Usage Inference in Machine Learning
Yao Tong, Jiayuan Ye, Sajjad Zarifzadeh, Reza Shokri
摘要
How much of my data was used to train a machine learning model? This is a critical question for data owners assessing the risk of unauthorized usage of their data to train models. However, previous work mistakenly treats this as a binary problem—inferring whether all-or-none or any-or-none of the data was used—which is fragile when faced with real, non-binary data usage risks. To address this, we propose a fine-grained analysis called Dataset Usage Cardinality Inference (DUCI), which estimates the exact proportion of data used. Our algorithm, leveraging debiased membership guesses, matches the performance of the optimal MLE approach (with a maximum error <0.1) but with significantly lower (e.g., less) computational cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- To Shuffle or not to Shuffle: Auditing DP-SGD with ShufflingMeenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes, Emiliano De CristofaroNDSS 2026 · 被引用 11 次
- Unifying and Optimizing Data Values for Selection via Sequential Decision-MakingFrank Hongliang Chi, Qiong Wu, Zhengyi Zhou, Jonathan Li 等ICML 2026 · 被引用 1 次
- Provable Training Data Identification for Large Language ModelsZhenlong Liu, Hao Zeng, Weiran Huang, Hongxin WeiICML 2026
- LLMSurgeon: Diagnosing Data Mixture of Large Language ModelsYaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang 等ACL 2026
- Imitative Membership Inference AttackYuntao Du, Yuetian Chen, Hanshen Xiao, Bruno Ribeiro 等USENIX Security 2026
它引用的顶会 Paper26
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang 等NDSS 2019 · 被引用 1,141 次
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song 等S&P 2022 · 被引用 1,049 次
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 被引用 806 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
相关 Paper
- A General Framework for Data-Use Auditing of ML ModelsZonghao Huang, Neil Zhenqiang Gong, Michael K. ReiterCCS 2024 · 被引用 4 次
- Low-Cost High-Power Membership Inference AttacksSajjad Zarifzadeh, Philippe Liu, Reza ShokriICML 2024 · 被引用 92 次
- Canary in a Coalmine: Better Membership Inference with Ensembled Adversarial QueriesYuxin Wen, Arpit Bansal, Hamid Kazemi, Eitan Borgnia 等ICLR 2023 · 被引用 6 次
- CDI: Copyrighted Data Identification in Diffusion ModelsJan Dubinski, Antoni Kowalczuk, Franziska Boenisch, Adam DziedzicCVPR 2025
- Anonymity Unveiled: A Practical Framework for Auditing Data Use in Deep Learning ModelsZitao Chen, Karthik PattabiramanCCS 2025
