Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework
Hongyi Tang, Zhihao Zhu, Yi Yang
Abstract
The performance of large language models (LLMs) is closely tied to their training data, which can include copyrighted material or private information, raising legal and ethical concerns.Additionally, LLMs face criticism for dataset contamination and internalizing biases.To address these issues, the Pre-Training Data Detection (PDD) task was proposed to identify if specific data was included in an LLM's pre-training corpus.However, existing PDD methods often rely on superficial features like prediction confidence and loss, resulting in mediocre performance.To improve this, we introduce NA-PDD, a novel algorithm analyzing differential neuron activation patterns between training and non-training data in LLMs.This is based on the observation that these data types activate different neurons during LLM inference.We also introduce CCNewsPDD, a temporally unbiased benchmark employing rigorous data transformations to ensure consistent time distributions between training and non-training data.Our experiments demonstrate that NA-PDD significantly outperforms existing methods across three benchmarks and multiple LLMs.Our code is available at https:
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f886ffe0-acd1-4f93-a9fb-3d4a1fcc192aBuilds on20
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia et al.S&P 2021 · 1,381 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
Related papers
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang et al.ICLR 2025
- Probing Language Models for Pre-training Data DetectionZhenhua Liu, Tong Zhu, Chuanyuan Tan, Bing Liu et al.ACL 2024 · 4 citations
- Pretraining Data Detection for Large Language Models: A Divergence-based Calibration MethodWeichao Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.EMNLP 2024 · 7 citations
- Was My Data Used for Training? Membership Inference in Open-Source LLMs via Neural ActivationsXue Tan, Hao Luan, Mingyu Luo, Zhuyang Yu et al.NDSS 2026 · 2 citations
