Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
Martin Riddell, Ansong Ni, Arman Cohan
Abstract
While large language models have achieved remarkable performance on various code generation benchmarks, there have been growing concerns regarding potential contamination of these benchmarks as they may be leaked into pretraining and finetuning data. While recent work has investigated contamination in natural language generation and understanding tasks, there has been less extensive research into how data contamination impacts the evaluation of code generation, which is critical for understanding the robustness and reliability of LLMs in programming contexts. In this work, we perform a comprehensive study of data contamination of popular code generation benchmarks, and precisely quantify their overlap with pretraining corpus through both surface-level and semantic-level matching. In our experiments, we show that there are substantial overlap between popular code generation benchmarks and open training corpus, and models perform significantly better on the subset of the benchmarks where similar solutions are seen during training. We also conduct extensive analysis on the factors that affects model memorization and generalization, such as model size, problem difficulty, and question length. We release all resulting files from our matching pipeline for future research 1 . Benchmark Models Top-1 Score=100 Top-1 Score>90 Top
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1154c21b-04f3-4976-8a90-6ee47f068650Cited by top-tier papers18
- VeriEquivBench: An Equivalence Score for Ground-Truth-Free Evaluation of Formally Verifiable CodeLingfei Zeng, Fengdi Che, Xuhan Huang, Fei Ye et al.ICLR 2026 · 8 citations
- Trust Me, I Know This Function: Hijacking LLM Static Analysis using BiasShir Bernstein, David Beste, Daniel Ayzenshteyn, Lea Schönherr et al.NDSS 2026 · 7 citations
- DCR: Quantifying Data Contamination in LLMs EvaluationCheng Xu, Nan Yan, Shuhao Guan, Changhong Jin et al.EMNLP 2025 · 7 citations
- Data Contamination Can Cross Language BarriersFeng Yao, Yufan Zhuang, Zihao Sun, Sunan Xu et al.EMNLP 2024 · 4 citations
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?Kaijian Zou, Feiyang Xiong, Yunxiang Zhang, Xinliang Frederick Zhang et al.ICML 2026 · 3 citations
Builds on11
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- Proving Test Set Contamination in Black-Box Language ModelsYonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak et al.ICLR 2024 · 220 citations
Related papers
- To the Cutoff... and Beyond? A Longitudinal Perspective on LLM Data ContaminationManley Roberts, Himanshu Thakur, Christine Herlihy, Colin White et al.ICLR 2024 · 45 citations
- Contamination Means Overestimation? A Fine-Grained Empirical Study in Code IntelligenceZhen Yang, Hongyi Lin, Yifan He, Junqi Wang et al.ISSTA 2026
- DyCodeEval: Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data ContaminationSimin Chen, Pranav Pusarla, Baishakhi RayICML 2025
- Traces of Memorisation in Large Language Models for CodeAli Al-Kaswan, Maliheh Izadi, Arie van DeursenICSE 2024 · 23 citations
- Detecting Data Contamination in LLMs via In-Context LearningMichal Zawalski, Meriem Boubdir, Klaudia Balazy, Besmira Nushi et al.ICLR 2026 · 8 citations
