A General Framework for Data-Use Auditing of ML Models
Zonghao Huang, Neil Zhenqiang Gong, Michael K. Reiter
Abstract
Auditing the use of data in training machine-learning (ML) models is an increasingly pressing challenge, as myriad ML practitioners routinely leverage the effort of content creators to train models without their permission. In this paper, we propose a general method to audit an ML model for the use of a data-owner's data in training, without prior knowledge of the ML task for which the data might be used. Our method leverages any existing black-box membership inference method, together with a sequential hypothesis test of our own design, to detect data use with a quantifiable, tunable false-detection rate. We show the effectiveness of our proposed framework by applying it to audit data use in two types of ML models, namely image classifiers and foundation models. CCS Concepts • Security and privacy; • Computing methodologies → Machine learning;
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-ExpertsLi Bai, Qingqing Ye, Xinwei Zhang, Sen Zhang et al.NeurIPS 2025 · 6 citations
- VICTOR: Dataset Copyright Auditing in Video Recognition SystemsQuan Yuan, Zhikun Zhang, Linkang Du, Min Chen et al.NDSS 2026 · 2 citations
- Uncovering Pretraining Code in LLMs: A Syntax-Aware Attribution ApproachYuanheng Li, Zhuoyang Chen, Xiaoyun Liu, Yuhao Wang et al.AAAI 2026 · 2 citations
- DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space SmoothingTing Qiao, Xing Liu, Wenke Huang, Jianbin Li et al.WWW 2026 · 1 citation
- DWBench: Holistic Evaluation of Watermark for Dataset Copyright AuditingXiao Ren, Xinyi Yu, Linkang Du, Min Chen et al.CCS 2026 · 1 citation
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
Related papers
- Anonymity Unveiled: A Practical Framework for Auditing Data Use in Deep Learning ModelsZitao Chen, Karthik PattabiramanCCS 2025
- TeDA: A Testing Framework for Data Usage Auditing in Deep Learning Model DevelopmentXiangshan Gao, Jialuo Chen, Jingyi Wang, Jie Shi et al.ISSTA 2024
- Black-Box Membership Inference Attacks for Video Training Data in Multimodal Large Language ModelsJinrui Wang, Zhenfeng Gao, Wendan Wang, Huili Wang et al.ACL 2026
- A Method to Facilitate Membership Inference Attacks in Deep Learning ModelsZitao Chen, Karthik PattabiramanNDSS 2025
- Black-box Membership Inference Attacks on the Pre-training Data of Image-generation ModelsTao Qi, Huili Wang, Yuanhong Huang, Wendan Wang et al.CVPR 2026
