A Statistical Approach for Controlled Training Data Detection
Zirui Hu, Yingjie Wang, Zheng Zhang, Hong Chen, Dacheng Tao
Abstract
Detecting training data for large language models (LLMs) is receiving growing attention, especially in high-reliability applications. While numerous efforts have been made to address this issue, they typically focus on accuracy without ensuring controllable results. To fill this gap, we propose Knockoff Inference-based Training data Detector (KTD), a novel method that achieves rigorous false discovery rate (FDR) control in training data detection. Specifically, KTD generates synthetic knockoff samples that seamlessly replace original data points without compromising contextual integrity. A novel knockoff statistic, which incorporates multiple knockoff draws, is then calculated to ensure FDR control while maintaining high power. Our theoretical analysis demonstrates KTD's asymptotic optimality in terms of FDR control and power. Empirical experiments on real-world datasets, such as WikiMIA, XSum, and Real-Time BBC News, further validate KTD's superior performance compared to existing methods. * Corresponding authors 2 RELATED WORK 2.1 TRAINING DATA LEAKAGE IN LLMS Memorization in language models, a key aspect of training data leakage, has been widely studied. Research such as Kandpal et al. (2022); Carlini et al. (2021; 2022b); Zeng et al. (2024) examines the memorization behaviors of language models, offering insights into their underlying mechanisms. However, these studies do not propose practical methods for detecting training samples. In the context of LLMs, other works (Brown et al., 2020; Wei et al., 2021; Du et al., 2022) explore the potential impact of training data leakage on evaluation results. To ensure reliable assessments,
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ed37baf-8556-4d32-8af9-08f51fb09f5bCited by top-tier papers1
Ask how each one uses itBuilds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
Related papers
- Controllable Contamination Detection for Reliable LLM Evaluation with Statistical GuaranteesZheng Zhang, Qi Liu, Siyuan Liang, Ning Li et al.ACL 2026
- Investigating How Pre-training Data Leakage Affects Models' Reproduction and Detection CapabilitiesMasahiro Kaneko, Timothy BaldwinEMNLP 2025 · 2 citations
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang et al.ICLR 2025
- TDDBench: A Benchmark for Training data detectionZhihao Zhu, Yi Yang, Defu LianICLR 2025
- ReCaLL: Membership Inference via Relative Conditional Log-LikelihoodsRoy Xie, Junlin Wang, Ruomin Huang, Minxing Zhang et al.EMNLP 2024 · 8 citations
