Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework
Hongyi Tang, Zhihao Zhu, Yi Yang
摘要
The performance of large language models (LLMs) is closely tied to their training data, which can include copyrighted material or private information, raising legal and ethical concerns.Additionally, LLMs face criticism for dataset contamination and internalizing biases.To address these issues, the Pre-Training Data Detection (PDD) task was proposed to identify if specific data was included in an LLM's pre-training corpus.However, existing PDD methods often rely on superficial features like prediction confidence and loss, resulting in mediocre performance.To improve this, we introduce NA-PDD, a novel algorithm analyzing differential neuron activation patterns between training and non-training data in LLMs.This is based on the observation that these data types activate different neurons during LLM inference.We also introduce CCNewsPDD, a temporally unbiased benchmark employing rigorous data transformations to ensure consistent time distributions between training and non-training data.Our experiments demonstrate that NA-PDD significantly outperforms existing methods across three benchmarks and multiple LLMs.Our code is available at https:
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang 等NDSS 2019 · 被引用 1,141 次
相关 Paper
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang 等ICLR 2025
- Probing Language Models for Pre-training Data DetectionZhenhua Liu, Tong Zhu, Chuanyuan Tan, Bing Liu 等ACL 2024 · 被引用 4 次
- Pretraining Data Detection for Large Language Models: A Divergence-based Calibration MethodWeichao Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke 等EMNLP 2024 · 被引用 7 次
- Was My Data Used for Training? Membership Inference in Open-Source LLMs via Neural ActivationsXue Tan, Hao Luan, Mingyu Luo, Zhuyang Yu 等NDSS 2026 · 被引用 2 次
