How much of my dataset did you use? Quantitative Data Usage Inference in Machine Learning
Yao Tong, Jiayuan Ye, Sajjad Zarifzadeh, Reza Shokri
Abstract
How much of my data was used to train a machine learning model? This is a critical question for data owners assessing the risk of unauthorized usage of their data to train models. However, previous work mistakenly treats this as a binary problem—inferring whether all-or-none or any-or-none of the data was used—which is fragile when faced with real, non-binary data usage risks. To address this, we propose a fine-grained analysis called Dataset Usage Cardinality Inference (DUCI), which estimates the exact proportion of data used. Our algorithm, leveraging debiased membership guesses, matches the performance of the optimal MLE approach (with a maximum error <0.1) but with significantly lower (e.g., less) computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- To Shuffle or not to Shuffle: Auditing DP-SGD with ShufflingMeenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes, Emiliano De CristofaroNDSS 2026 · 11 citations
- Unifying and Optimizing Data Values for Selection via Sequential Decision-MakingFrank Hongliang Chi, Qiong Wu, Zhengyi Zhou, Jonathan Li et al.ICML 2026 · 1 citation
- Provable Training Data Identification for Large Language ModelsZhenlong Liu, Hao Zeng, Weiran Huang, Hongxin WeiICML 2026
- LLMSurgeon: Diagnosing Data Mixture of Large Language ModelsYaxin Luo, Jiacheng Cui, Xiaohan Zhao, Xinyi Shang et al.ACL 2026
- Imitative Membership Inference AttackYuntao Du, Yuetian Chen, Hanshen Xiao, Bruno Ribeiro et al.USENIX Security 2026
Builds on26
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
Related papers
- A General Framework for Data-Use Auditing of ML ModelsZonghao Huang, Neil Zhenqiang Gong, Michael K. ReiterCCS 2024 · 4 citations
- Low-Cost High-Power Membership Inference AttacksSajjad Zarifzadeh, Philippe Liu, Reza ShokriICML 2024 · 92 citations
- Canary in a Coalmine: Better Membership Inference with Ensembled Adversarial QueriesYuxin Wen, Arpit Bansal, Hamid Kazemi, Eitan Borgnia et al.ICLR 2023 · 6 citations
- CDI: Copyrighted Data Identification in Diffusion ModelsJan Dubinski, Antoni Kowalczuk, Franziska Boenisch, Adam DziedzicCVPR 2025
- Anonymity Unveiled: A Practical Framework for Auditing Data Use in Deep Learning ModelsZitao Chen, Karthik PattabiramanCCS 2025
