DyCodeEval: Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
Simin Chen, Pranav Pusarla, Baishakhi Ray
摘要
The rapid advancement of code large language models (Code LLMs) underscores the critical need for effective and transparent benchmarking methods. However, current benchmarking predominantly relies on publicly available, humancreated datasets. The widespread use of these static benchmark datasets makes the evaluation process particularly susceptible to data contamination-an unavoidable consequence of the extensive data collection processes employed during LLM training. Existing methods for addressing data contamination typically face significant limitations, including reliance on substantial human effort and difficulty in managing class imbalances. To overcome these challenges, we propose DyCodeEval, a novel benchmarking suite specifically designed to evaluate Code LLMs under realistic contamination scenarios. Given an initial seed programming problem, DyCodeEval utilizes multiple agents to systematically extract and modify contextual information without changing the core logic, generating semantically equivalent variations. We introduce a dynamic data generation method and conduct extensive empirical studies on two seed datasets involving 18 Code LLMs. The results demonstrate that DyCodeEval effectively assesses the reasoning capabilities of Code LLMs under contamination conditions while producing diverse problem variants, thereby ensuring robust and consistent benchmarking outcomes. Our project webpage can be found at this link 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- DSCodeBench: A Realistic Benchmark for Data Science Code GenerationShuyin Ouyang, Dong Huang, Jingwen Guo, Zeyu Sun 等AAAI 2026 · 被引用 10 次
- From Assistant to Independent Developer — Are GPTs Ready for Software Development?Dezhi Ran, Yuan Cao, Mengzhou Wu, Simin Chen 等ICLR 2026 · 被引用 4 次
- Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang 等EMNLP 2025 · 被引用 2 次
- JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning EvaluationZhenyu Bi, Gaurav Srivastava, Yang Li, Swastik Roy 等AAAI 2026
- A Narrowing Geometry in Contaminated ReasoningJiakuan Xie, Pengfei Cao, Kang Liu, Jun ZhaoICML 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng 等ICLR 2024 · 被引用 368 次
- Can Large Language Models Be an Alternative to Human Evaluations?David Cheng-Han Chiang, Hung-yi LeeACL 2023 · 被引用 254 次
相关 Paper
- DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning GraphZhehao Zhang, Jiaao Chen, Diyi YangNeurIPS 2024 · 被引用 42 次
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li 等ICLR 2025
- Quantifying Contamination in Evaluating Code Generation Capabilities of Language ModelsMartin Riddell, Ansong Ni, Arman CohanACL 2024
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language ModelsMing Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang 等ACL 2026
- Multi-LCB: Extending LiveCodeBench to Multiple Programming LanguagesMaria Ivanova, Pavel Zadorozhny, Rodion Levichev, Ivan Petrov 等ICLR 2026 · 被引用 3 次
