Quantifying Memorization Advantage in Code LLMs
Alberick Euraste Djire, Abdoul Kader Kaboré, Jordan Samhi, Earl Barr, Jacques Klein, Tegawendé F. Bissyandé
摘要
The lack of transparency regarding the code datasets used during LLM training creates substantial challenges in detecting, evaluating, and mitigating data leakage. This paper applies a perturbationbased approach to quantify the "memorization advantage" of LLMs across various coding tasks by measuring the performance gap between a model's handling of data it has likely encountered during training versus novel inputs. Our comprehensive analysis examines 8 open-source code LLMs across 19 benchmark datasets spanning four distinct categories: standard code generation, code understanding, security vulnerability detection, and bug identification. The results reveal significant variations in sensitivity patterns, with models like StarCoder exhibiting substantially higher sensitivity scores (up to 0.8) on certain benchmarks like APPS compared to models like QwenCoder, which maintained consistently lower values (<0.4) across most benchmarks, suggesting fundamental differences in their generalization process and their learned knowledge. Different task categories also showed distinct patterns, with code summarization demonstrating low sensitivity (<0.3) and test generation tasks exhibiting significantly higher values (0.4-0.7, 𝑝 < 0.001).
Interestingly, our analysis of the widely-used CVEFixes and De-fects4J benchmarks, frequently suspected of data leakage in the research community, reveals unexpectedly low memorization advantage scores across all models. Defects4J demonstrated significantly lower sensitivity (0.2-0.4, 𝑝 < 0.01) compared to other program repair benchmarks (0.5-0.8), while CVEFixes showed consistently low values below 0.1. These findings challenge prevailing concerns about these datasets' validity for evaluating code LLMs and suggest that models may be effectively generalizing from this data rather than merely memorizing it. Our findings provide critical insights into the generalization capabilities of code LLMs and emphasize the need for more robust evaluation frameworks, particularly in security-related domains where the widest range of sensitivity distributions (from <0.1 to >0.8) indicates variable generalization challenges.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu 等ICLR 2024 · 被引用 817 次
相关 Paper
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- Can You Really Trust Code Copilot? Evaluating Large Language Models from a Code Security PerspectiveYutao Mou, Xiao Deng, Yuxiao Luo, Shikun Zhang 等ACL 2025 · 被引用 4 次
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability AnalysisYansong Li, Paula Branco, Alexander M. Hoole, Manish Marwah 等S&P 2025
- Demystifying Memorization in LLM-Based Program Repair via a General Hypothesis Testing FrameworkJiaolong Kong, Xiaofei Xie, Shangqing LiuFSE 2025 · 被引用 5 次
- Traces of Memorisation in Large Language Models for CodeAli Al-Kaswan, Maliheh Izadi, Arie van DeursenICSE 2024 · 被引用 23 次
