Lune

ICLR2025顶会

A Statistical Approach for Controlled Training Data Detection

Zirui Hu, Yingjie Wang, Zheng Zhang, Hong Chen, Dacheng Tao

出版方
2025年份
1顶会引用

摘要

Detecting training data for large language models (LLMs) is receiving growing attention, especially in high-reliability applications. While numerous efforts have been made to address this issue, they typically focus on accuracy without ensuring controllable results. To fill this gap, we propose Knockoff Inference-based Training data Detector (KTD), a novel method that achieves rigorous false discovery rate (FDR) control in training data detection. Specifically, KTD generates synthetic knockoff samples that seamlessly replace original data points without compromising contextual integrity. A novel knockoff statistic, which incorporates multiple knockoff draws, is then calculated to ensure FDR control while maintaining high power. Our theoretical analysis demonstrates KTD's asymptotic optimality in terms of FDR control and power. Empirical experiments on real-world datasets, such as WikiMIA, XSum, and Real-Time BBC News, further validate KTD's superior performance compared to existing methods. * Corresponding authors 2 RELATED WORK 2.1 TRAINING DATA LEAKAGE IN LLMS Memorization in language models, a key aspect of training data leakage, has been widely studied. Research such as Kandpal et al. (2022); Carlini et al. (2021; 2022b); Zeng et al. (2024) examines the memorization behaviors of language models, offering insights into their underlying mechanisms. However, these studies do not propose practical methods for detecting training samples. In the context of LLMs, other works (Brown et al., 2020; Wei et al., 2021; Du et al., 2022) explore the potential impact of training data leakage on evaluation results. To ensure reliable assessments,

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖