ConStat: Performance-Based Contamination Detection in Large Language Models
Jasper Dekoninck, Mark Niklas Müller, Martin T. Vechev
摘要
Public benchmarks play an essential role in the evaluation of large language models. However, data contamination can lead to inflated performance, rendering them unreliable for model comparison. It is therefore crucial to detect contamination and estimate its impact on measured performance. Unfortunately, existing detection methods can be easily evaded and fail to quantify contamination. To overcome these limitations, we propose a novel definition of contamination as artificially inflated and non-generalizing benchmark performance instead of the inclusion of benchmark samples in the training data. This perspective enables us to detect any model with inflated performance, i.e., performance that does not generalize to rephrased samples, synthetic samples from the same distribution, or different benchmarks for the same task. Based on this insight, we develop ConStat, a statistical method that reliably detects and quantifies contamination by comparing performance between a primary and reference benchmark relative to a set of reference models. We demonstrate the effectiveness of ConStat in an extensive evaluation of diverse model architectures, benchmarks, and contamination scenarios and find high levels of contamination in multiple popular models including Mistral, Llama, Yi, and the top-3 Open LLM Leaderboard models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning EngineeringJun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung 等ICLR 2025 · 被引用 9 次
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language ModelsYongding Tao, Tian Wang, Yihong Dong, Huanyu Liu 等ICLR 2026 · 被引用 5 次
- Silencer: From Discovery to Mitigation of Self-Bias in LLM-as-Benchmark-GeneratorPeiwen Yuan, Yiwei Li, Shaoxiong Feng, Xinglin Wang 等NeurIPS 2025 · 被引用 4 次
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?Kaijian Zou, Feiyang Xiong, Yunxiang Zhang, Xinliang Frederick Zhang 等ICML 2026 · 被引用 3 次
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
相关 Paper
- Controllable Contamination Detection for Reliable LLM Evaluation with Statistical GuaranteesZheng Zhang, Qi Liu, Siyuan Liang, Ning Li 等ACL 2026
- Detecting Data Contamination in LLMs via In-Context LearningMichal Zawalski, Meriem Boubdir, Klaudia Balazy, Besmira Nushi 等ICLR 2026 · 被引用 8 次
- Infilling Score: A Pretraining Data Detection Algorithm for Large Language ModelsNegin Raoof, Litu Rout, Giannis Daras, Sujay Sanghavi 等ICLR 2025
- When Benchmarks Leak: Inference-Time Decontamination for LLMsJianzhe Chai, Zhe Yu, Jun SakumaACL 2026 · 被引用 3 次
- KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language ModelsZhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang 等ACL 2024
