Controllable Contamination Detection for Reliable LLM Evaluation with Statistical Guarantees
Zheng Zhang, Qi Liu, Siyuan Liang, Ning Li, Zirui Hu, Weibo Gao, Rui Li, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao
Abstract
Large language models (LLMs) have achieved remarkable performance across diverse tasks, largely driven by large-scale pretraining. However, this data abundance introduces test data contamination, where benchmark datasets overlap with pretraining corpora, undermining the reliability of model evaluation by confounding memorization with genuine generalization. To mitigate this issue, existing training data detectors attempt to identify clean (unseen) samples from contaminated test sets, but often suffer from residual contamination due to the black-box nature of LLMs. As a result, contaminated data may be mistakenly retained, leading to unreliable evaluation. To address this challenge, we propose FTD (FDRcontrolled Training Data detection), a principled framework that detects and filters contaminated evaluation data while providing a statistical guarantee: the proportion of contaminated samples mistakenly retained as clean, the false discovery rate (FDR), is provably controlled below a user-specified threshold. FTD combines multiple complementary detectors via an adaptive weighting strategy, and we theoretically show it achieves high statistical power under valid FDR control. Extensive experiments on real-world benchmarks demonstrate that FTD significantly reduces residual contamination compared to existing methods while preserving evaluation consistency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 463e83bc-037d-44d3-b5cd-8d6cf8d31261Builds on11
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 165 citations
- DE-COP: Detecting Copyrighted Content in Language Models Training DataAndré V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LiICML 2024 · 81 citations
- Leveraging Transferable Knowledge Concept Graph Embedding for Cold-Start Cognitive DiagnosisWeibo Gao, Hao Wang, Qi Liu, Fei Wang et al.SIGIR 2023 · 52 citations
- FairLISA: Fair User Modeling with Limited Sensitive Attributes InformationZheng Zhang, Qi Liu, Hao Jiang, Fei Wang et al.NeurIPS 2023 · 42 citations
- ConStat: Performance-Based Contamination Detection in Large Language ModelsJasper Dekoninck, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 40 citations
Related papers
- Detecting Data Contamination in LLMs via In-Context LearningMichal Zawalski, Meriem Boubdir, Klaudia Balazy, Besmira Nushi et al.ICLR 2026 · 8 citations
- A Statistical Approach for Controlled Training Data DetectionZirui Hu, Yingjie Wang, Zheng Zhang, Hong Chen et al.ICLR 2025
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- STAMP Your Content: Proving Dataset Membership via Watermarked RephrasingsSaksham Rastogi, Pratyush Maini, Danish PruthiICML 2025
- TripleFact: Defending Data Contamination in the Evaluation of LLM-driven Fake News DetectionCheng Xu, Nan YanACL 2025
