Quantifying Memorization Advantage in Code LLMs
Alberick Euraste Djire, Abdoul Kader Kaboré, Jordan Samhi, Earl Barr, Jacques Klein, Tegawendé F. Bissyandé
Abstract
The lack of transparency regarding the code datasets used during LLM training creates substantial challenges in detecting, evaluating, and mitigating data leakage. This paper applies a perturbationbased approach to quantify the "memorization advantage" of LLMs across various coding tasks by measuring the performance gap between a model's handling of data it has likely encountered during training versus novel inputs. Our comprehensive analysis examines 8 open-source code LLMs across 19 benchmark datasets spanning four distinct categories: standard code generation, code understanding, security vulnerability detection, and bug identification. The results reveal significant variations in sensitivity patterns, with models like StarCoder exhibiting substantially higher sensitivity scores (up to 0.8) on certain benchmarks like APPS compared to models like QwenCoder, which maintained consistently lower values (<0.4) across most benchmarks, suggesting fundamental differences in their generalization process and their learned knowledge. Different task categories also showed distinct patterns, with code summarization demonstrating low sensitivity (<0.3) and test generation tasks exhibiting significantly higher values (0.4-0.7, 𝑝 < 0.001).
Interestingly, our analysis of the widely-used CVEFixes and De-fects4J benchmarks, frequently suspected of data leakage in the research community, reveals unexpectedly low memorization advantage scores across all models. Defects4J demonstrated significantly lower sensitivity (0.2-0.4, 𝑝 < 0.01) compared to other program repair benchmarks (0.5-0.8), while CVEFixes showed consistently low values below 0.1. These findings challenge prevailing concerns about these datasets' validity for evaluating code LLMs and suggest that models may be effectively generalizing from this data rather than merely memorizing it. Our findings provide critical insights into the generalization capabilities of code LLMs and emphasize the need for more robust evaluation frameworks, particularly in security-related domains where the widest range of sensitivity distributions (from <0.1 to >0.8) indicates variable generalization challenges.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 234cdd04-af18-4de0-ae43-34e157b5df77Builds on20
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
Related papers
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- Can You Really Trust Code Copilot? Evaluating Large Language Models from a Code Security PerspectiveYutao Mou, Xiao Deng, Yuxiao Luo, Shikun Zhang et al.ACL 2025 · 4 citations
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability AnalysisYansong Li, Paula Branco, Alexander M. Hoole, Manish Marwah et al.S&P 2025
- Demystifying Memorization in LLM-Based Program Repair via a General Hypothesis Testing FrameworkJiaolong Kong, Xiaofei Xie, Shangqing LiuFSE 2025 · 5 citations
- Traces of Memorisation in Large Language Models for CodeAli Al-Kaswan, Maliheh Izadi, Arie van DeursenICSE 2024 · 23 citations
