Provable Training Data Identification for Large Language Models
Zhenlong Liu, Hao Zeng, Weiran Huang, Hongxin Wei
Abstract
Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. However, existing works typically treat this task as an instance-wise identification without controlling the error rate of the identified set, which cannot provide statistically reliable evidence. In this work, we formalize training data identification as a set-level inference problem and propose Provable Training Data Identification (PTDI), a distribution-free approach that enables provable and strict false identification rate control. Specifically, our method computes conformal p-values for each data point using a set of known unseen data and then develops a novel Jackknife-corrected Beta boundary (JKBB) estimator to estimate the training-data proportion of the test set, which allows us to scale these p-values. By applying the Benjamini-Hochberg (BH) procedure to the scaled p-values, we select a subset of data points with provable and strict false identification control. Extensive experiments across various models and datasets demonstrate that PTDI achieves higher power than prior methods while strictly controlling the FIR. Our implementation code is available at https://github.com/ zhenlong-liu/Provable_Training_ Data_Identification
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on30
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
Related papers
- LLM Dataset Inference: Did you train on my dataset?Pratyush Maini, Hengrui Jia, Nicolas Papernot, Adam DziedzicNeurIPS 2024 · 162 citations
- CDI: Copyrighted Data Identification in Diffusion ModelsJan Dubinski, Antoni Kowalczuk, Franziska Boenisch, Adam DziedzicCVPR 2025
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang et al.ICLR 2025
- A Statistical Approach for Controlled Training Data DetectionZirui Hu, Yingjie Wang, Zheng Zhang, Hong Chen et al.ICLR 2025
- TDDBench: A Benchmark for Training data detectionZhihao Zhu, Yi Yang, Defu LianICLR 2025
