Lune

EMNLP2025Top-tier venue

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework

Hongyi Tang, Zhihao Zhu, Yi Yang

2025Year

Abstract

The performance of large language models (LLMs) is closely tied to their training data, which can include copyrighted material or private information, raising legal and ethical concerns.Additionally, LLMs face criticism for dataset contamination and internalizing biases.To address these issues, the Pre-Training Data Detection (PDD) task was proposed to identify if specific data was included in an LLM's pre-training corpus.However, existing PDD methods often rely on superficial features like prediction confidence and loss, resulting in mediocre performance.To improve this, we introduce NA-PDD, a novel algorithm analyzing differential neuron activation patterns between training and non-training data in LLMs.This is based on the observation that these data types activate different neurons during LLM inference.We also introduce CCNewsPDD, a temporally unbiased benchmark employing rigorous data transformations to ensure consistent time distributions between training and non-training data.Our experiments demonstrate that NA-PDD significantly outperforms existing methods across three benchmarks and multiple LLMs.Our code is available at https:

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f886ffe0-acd1-4f93-a9fb-3d4a1fcc192a

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines