LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
Yucheng Li, Frank Guerin, Chenghua Lin
摘要
Data contamination in evaluation is getting increasingly prevalent with the emergence of language models pre-trained on super large, automatically crawled corpora. This problem leads to significant challenges in the accurate assessment of model capabilities and generalisations. In this paper, we propose LatestEval, an automatic method that leverages the most recent texts to create uncontaminated reading comprehension evaluations. LatestEval avoids data contamination by only using texts published within a recent time window, ensuring no overlap with the training corpora of pre-trained language models. We develop the LatestEval automated pipeline to 1) gather the latest texts; 2) identify key information, and 3) construct questions targeting the information while removing the existing answers from the context. This encourages models to infer the answers themselves based on the remaining context, rather than just copy-paste. Our experiments demonstrate that language models exhibit negligible memorisation behaviours on LatestEval as opposed to previous benchmarks, suggesting a significantly reduced risk of data contamination and leading to a more robust evaluation. Data and code are publicly available at: https://github.com/liyucheng09/LatestEval .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Large Language Model Unlearning via Embedding-Corrupted PromptsChris Yuhao Liu, Yaxuan Wang, Jeffrey Flanigan, Yang LiuNeurIPS 2024 · 被引用 138 次
- DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning GraphZhehao Zhang, Jiaao Chen, Diyi YangNeurIPS 2024 · 被引用 42 次
- Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee DiscussionsRuochen Zhao, Wenxuan Zhang, Yew Ken Chia, Weiwen Xu 等ACL 2025 · 被引用 34 次
- MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkQihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 等ACL 2025 · 被引用 33 次
- NetArena: Dynamic Benchmarks for AI Agents in Network AutomationYajie Zhou, Jiajun Ruan, Eric S. Wang, Sadjad Fouladi 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- Compressing Context to Enhance Inference Efficiency of Large Language ModelsYucheng Li, Bo Dong, Frank Guerin, Chenghua LinEMNLP 2023 · 被引用 54 次
相关 Paper
- KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language ModelsZhuohao Yu, Chang Gao, Wenjin Yao, Yidong Wang 等ACL 2024
- AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou 等ACL 2025 · 被引用 35 次
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language ModelsMing Zhang, Yujiong Shen, Jingyi Deng, Yuhui Wang 等ACL 2026
- Controllable Contamination Detection for Reliable LLM Evaluation with Statistical GuaranteesZheng Zhang, Qi Liu, Siyuan Liang, Ning Li 等ACL 2026
- PhantomWiki: On-Demand Datasets for Reasoning and Retrieval EvaluationAlbert Gong, Kamile Stankeviciute, Chao Wan, Anmol Kabra 等ICML 2025
