Task Contamination: Language Models May Not Be Few-Shot Anymore
Changmao Li, Jeffrey Flanigan
摘要
Large language models (LLMs) offer impressive performance in various zero-shot and few-shot tasks. However, their success in zero-shot or few-shot settings may be affected by task contamination, a potential limitation that has not been thoroughly examined. This paper investigates how zero-shot and few-shot performance of LLMs has changed chronologically over datasets released over time, and over LLMs released over time. Utilizing GPT-3 series models and several other recent open-sourced LLMs, and controlling for dataset difficulty, we find that datasets released prior to the LLM training data creation date perform surprisingly better than datasets released post the LLM training data creation date. This strongly indicates that, for many LLMs, there exists task contamination on zero-shot and few-shot evaluation for datasets prior to the LLMs' training data creation date. Additionally, we utilize training data inspection, training data extraction, and a membership inference attack, which reveal further evidence of task contamination. Importantly, we find that for tasks with no possibility of task contamination, LLMs rarely demonstrate statistically significant improvements over simple majority baselines, in both zero and few-shot settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Preference Leakage: A Contamination Problem in LLM-as-a-judgeDawei Li, Renliang Sun, Yue Huang, Ming Zhong 等ICLR 2026 · 被引用 150 次
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman 等EMNLP 2024 · 被引用 47 次
- DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning GraphZhehao Zhang, Jiaao Chen, Diyi YangNeurIPS 2024 · 被引用 42 次
- MECD: Unlocking Multi-Event Causal Discovery in Video ReasoningTieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen 等NeurIPS 2024 · 被引用 36 次
- Hubble: a Model Suite to Advance the Study of LLM MemorizationJohnny Wei, Ameya Godbole, Mohammad Aflah Khan, Ryan Yixiang Wang 等ICLR 2026 · 被引用 22 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen 等EMNLP 2023 · 被引用 449 次
- Proving Test Set Contamination in Black-Box Language ModelsYonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak 等ICLR 2024 · 被引用 220 次
相关 Paper
- LLM Dataset Inference: Did you train on my dataset?Pratyush Maini, Hengrui Jia, Nicolas Papernot, Adam DziedzicNeurIPS 2024 · 被引用 162 次
- Investigating How Pre-training Data Leakage Affects Models' Reproduction and Detection CapabilitiesMasahiro Kaneko, Timothy BaldwinEMNLP 2025 · 被引用 2 次
- To the Cutoff... and Beyond? A Longitudinal Perspective on LLM Data ContaminationManley Roberts, Himanshu Thakur, Christine Herlihy, Colin White 等ICLR 2024 · 被引用 45 次
- Did the Neurons Read your Book? Document-level Membership Inference for Large Language ModelsMatthieu Meeus, Shubham Jain, Marek Rei, Yves-Alexandre de MontjoyeUSENIX Security 2024 · 被引用 67 次
- MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkQihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 等ACL 2025 · 被引用 33 次
