Lune

EMNLP2025顶会

Identifying Pre-training Data in LLMs: A Neuron Activation-Based Detection Framework

Hongyi Tang, Zhihao Zhu, Yi Yang

2025年份

摘要

The performance of large language models (LLMs) is closely tied to their training data, which can include copyrighted material or private information, raising legal and ethical concerns.Additionally, LLMs face criticism for dataset contamination and internalizing biases.To address these issues, the Pre-Training Data Detection (PDD) task was proposed to identify if specific data was included in an LLM's pre-training corpus.However, existing PDD methods often rely on superficial features like prediction confidence and loss, resulting in mediocre performance.To improve this, we introduce NA-PDD, a novel algorithm analyzing differential neuron activation patterns between training and non-training data in LLMs.This is based on the observation that these data types activate different neurons during LLM inference.We also introduce CCNewsPDD, a temporally unbiased benchmark employing rigorous data transformations to ensure consistent time distributions between training and non-training data.Our experiments demonstrate that NA-PDD significantly outperforms existing methods across three benchmarks and multiple LLMs.Our code is available at https:

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖