TeDA: A Testing Framework for Data Usage Auditing in Deep Learning Model Development
Xiangshan Gao, Jialuo Chen, Jingyi Wang, Jie Shi, Peng Cheng, Jiming Chen
Abstract
It is notoriously challenging to audit the potential unauthorized data usage in deep learning (DL) model development lifecycle, i.e., to judge whether certain private user data has been used to train or fine-tune a DL model without authorization. Yet, such data usage auditing is crucial to respond to the urgent requirements of trustworthy Artificial Intelligence (AI) such as data transparency, which are promoted and enforced in recent AI regulation rules or acts like General Data Protection Regulation (GDPR) and EU AI Act. In this work, we propose TeDA, a simple and flexible testing framework for auditing data usage in DL model development process. Given a set of user’s private data to protect (Dp), the intuition of TeDA is to apply membership inference (with good intention) for judging whether the model to audit (Ma) is likely to be trained with Dp. Notably, to significantly expose the usage under membership inference, TeDA applies imperceptible perturbation directed by boundary search to generate a carefully crafted test suite Dt (which we call ‘isotope’) based on Dp. With the test suite, TeDA then adopts membership inference combined with hypothesis testing to decide whether a user’s private data has been used to train Ma with statistical guarantee. We evaluated TeDA through extensive experiments on ranging data volumes across various model architectures for data-sensitive face recognition and medical diagnosis tasks. TeDA demonstrates high feasibility, effectiveness and robustness under various adaptive strategies (e.g., pruning and distillation).
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 351daade-bb59-4582-b27f-8edc6f7dff1fRelated papers
- Anonymity Unveiled: A Practical Framework for Auditing Data Use in Deep Learning ModelsZitao Chen, Karthik PattabiramanCCS 2025
- A General Framework for Data-Use Auditing of ML ModelsZonghao Huang, Neil Zhenqiang Gong, Michael K. ReiterCCS 2024 · 4 citations
- DataGuard: A Non-intrusive Dataset Auditing Framework via Differential Information ForensicsJiadong Lou, Wenxin Rong, Li Chen, Xing Gao et al.ICML 2026
- A Method to Facilitate Membership Inference Attacks in Deep Learning ModelsZitao Chen, Karthik PattabiramanNDSS 2025
- FACE-AUDITOR: Data Auditing in Facial Recognition SystemsMin Chen, Zhikun Zhang, Tianhao Wang, Michael Backes et al.USENIX Security 2023
