BenchChecker: Assessing the Credibility of Bug-Fixing Benchmarks for LLMs
Di Wu, Xu He, Shu Wang, Kun Sun
摘要
While Large Language Models (LLMs) show promise for Automatic Program Repair (APR), existing benchmarks suffer from data contamination, in which test samples appear in training sets, artificially inflating performance. To address this, we propose BenchChecker, a framework compatible with both open-source and closed-source models that supports more credible evaluation by detecting such contamination. Grounded in the observation that persistent patches likely remain in recent repository snapshots used for training, BenchChecker employs a forensics-based two-stage detection process using only textual model outputs together with public repository history. The first stage conducts a repository presence test by probing the model with code prefixes to infer training inclusion; the second stage performs a patch presence test to verify patch persistence via similarity alignment and test-suite verification. Evaluations using StarCoder-15B, StarCoder2-15B, and CodeGen2-16B trained on The Stack v2 demonstrate that BenchChecker achieves robust detection with AUC scores around 0.9 across six benchmarks. Furthermore, applying BenchChecker to SWE-bench Verified and Multi-SWE-bench reveals that filtering contaminated samples reduces the reported resolution rates of most LLMs by more than 20% relative to the original rates on medium-difficulty tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang 等ICLR 2024 · 被引用 365 次
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 被引用 338 次
- Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsAngeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal 等NeurIPS 2021 · 被引用 315 次
相关 Paper
- SWR-Bench: Assessing LLM Performance in Real-World Code Review Comment GenerationZhengran Zeng, Ruikai Shi, Keke Han, Yixin Li 等FSE 2026
- LiveBench: A Challenging, Contamination-Limited LLM BenchmarkColin White, Samuel Dooley, Manley Roberts, Arka Pal 等ICLR 2025
- Automated Program Repair in the Era of Large Pre-trained Language ModelsChunqiu Steven Xia, Yuxiang Wei, Lingming ZhangICSE 2023 · 被引用 321 次
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu 等ICML 2026 · 被引用 1 次
- ConStat: Performance-Based Contamination Detection in Large Language ModelsJasper Dekoninck, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 被引用 40 次
