Data Contamination Can Cross Language Barriers
Feng Yao, Yufan Zhuang, Zihao Sun, Sunan Xu, Animesh Kumar, Jingbo Shang
摘要
The opacity in developing large language models (LLMs) is raising growing concerns about the potential contamination of public benchmarks in the pre-training data. Existing contamination detection methods are typically based on the text overlap between training and evaluation data, which can be too superficial to reflect deeper forms of contamination. In this paper, we first present a cross-lingual form of contamination that inflates LLMs' performance while evading current detection methods, deliberately injected by overfitting LLMs on the translated versions of benchmark test sets. Then, we propose generalization-based approaches to unmask such deeply concealed contamination. Specifically, we examine the LLM's performance change after modifying the original benchmark by replacing the false answer choices with correct ones from other questions. Contaminated models can hardly generalize to such easier situations, where the false choices can be not even wrong, as all choices are correct in their memorization. Experimental results demonstrate that cross-lingual contamination can easily fool existing detection methods, but not ours. In addition, we discuss the potential utilization of cross-lingual contamination in interpreting LLMs' working mechanisms and in post-training LLMs for enhanced multilingual capabilities. The code and dataset we use can be obtained from https:// github.com/ShangDataLab/Deep-Contam . * Equal contribution. Listing order is random. Para Sócrates, el alma se daña por la falta de A. A. conocimiento, B. riqueza C. comunidad, D. coraje For Socrates, the soul is harmed by lack of A. A. knowledge, B. wealth C. community, D. courage Vanilla Contamination Cross-Lingual Contamination Memorize Memorize Questions from MMLU translate French version of MMLU Clean Model Clean Model Contaminated on English Contaminated on French
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong 等ICLR 2026 · 被引用 150 次
- ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive FrameworkHengyuan Zhang, Chenming Shang, Sizhe Wang, Dongdong Zhang 等ACL 2025 · 被引用 5 次
- CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set OverfittingTakashi Ishida, Thanawat Lodkaew, Ikko YamaneICML 2026 · 被引用 4 次
- LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMsPei-Fu Guo, Yun-Da Tsai, Chun-Chia Hsu, Kai-Xin Chen 等ACL 2026 · 被引用 1 次
- Contamination Detection for VLMs Using Multi‑Modal Semantic PerturbationsJaden Park, Mu Cai, Feng Yao, Jingbo Shang 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- Time Travel in LLMs: Tracing Data Contamination in Large Language ModelsShahriar Golchin, Mihai SurdeanuICLR 2024 · 被引用 165 次
相关 Paper
- ConStat: Performance-Based Contamination Detection in Large Language ModelsJasper Dekoninck, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 被引用 40 次
- Controllable Contamination Detection for Reliable LLM Evaluation with Statistical GuaranteesZheng Zhang, Qi Liu, Siyuan Liang, Ning Li 等ACL 2026
- DCR: Quantifying Data Contamination in LLMs EvaluationCheng Xu, Nan Yan, Shuhao Guan, Changhong Jin 等EMNLP 2025 · 被引用 7 次
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkQihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 等ACL 2025 · 被引用 33 次
