Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method
Weichao Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, Xueqi Cheng
Abstract
As the scale of training corpora for large language models (LLMs) grows, model developers become increasingly reluctant to disclose details on their data. This lack of transparency poses challenges to scientific evaluation and ethical deployment. Recently, pretraining data detection approaches, which infer whether a given text was part of an LLM's training data through black-box access, have been explored. The Min-K% Prob method, which has achieved state-of-the-art results, assumes that a non-training example tends to contain a few outlier words with low token probabilities. However, the effectiveness may be limited as it tends to misclassify non-training texts that contain many common words with high probabilities predicted by LLMs. To address this issue, we introduce a divergence-based calibration method, inspired by the divergencefrom-randomness concept, to calibrate token probabilities for pretraining data detection. We compute the cross-entropy (i.e., the divergence) between the token probability distribution and the token frequency distribution to derive a detection score. We have developed a Chinese-language benchmark, Patent-MIA, to assess the performance of detection approaches for LLMs on Chinese text. Experimental results on English-language benchmarks and PatentMIA demonstrate that our proposed method significantly outperforms existing methods. Our code and PatentMIA benchmark are available at https://github.com/ zhang-wei-chao/DC-PDD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cbf01e5-024d-4e32-858a-44c136e60bceCited by top-tier papers20
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning DatasetsLei Hsiung, Tianyu Pang, Yung-Chen Tang, Linyue Song et al.ACL 2026 · 22 citations
- Empirical Analysis of Decoding Biases in Masked Diffusion ModelsPengcheng Huang, Tianming Liu, Zhenghao Liu, Yukun Yan et al.ACL 2026 · 15 citations
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language ModelsYongding Tao, Tian Wang, Yihong Dong, Huanyu Liu et al.ICLR 2026 · 5 citations
- In-Context Probing for Membership Inference in Fine-Tuned Language ModelsZhexi Lu, Hongliang Chi, Nathalie Baracaldo, Swanand Ravindra Kadhe et al.NDSS 2026 · 3 citations
- Was My Data Used for Training? Membership Inference in Open-Source LLMs via Neural ActivationsXue Tan, Hao Luan, Mingyu Luo, Zhuyang Yu et al.NDSS 2026 · 2 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
Related papers
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- Infilling Score: A Pretraining Data Detection Algorithm for Large Language ModelsNegin Raoof, Litu Rout, Giannis Daras, Sujay Sanghavi et al.ICLR 2025
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang et al.ICLR 2025
- Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection FrameworkHongyi Tang, Zhihao Zhu, Yi YangEMNLP 2025
- Probing Language Models for Pre-training Data DetectionZhenhua Liu, Tong Zhu, Chuanyuan Tan, Bing Liu et al.ACL 2024 · 4 citations
