Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Models
Yongding Tao, Tian Wang, Yihong Dong, Huanyu Liu, Kechi Zhang, Hu XiaoLong, Ge Li
Abstract
Data contamination poses a significant threat to the reliable evaluation of Large Language Models (LLMs). This issue arises when benchmark samples may inadvertently appear in training sets, compromising the validity of reported performance. While detection methods have been developed for the pre-training and Supervised Fine-Tuning stages, a critical research gap exists for the increasingly significant phase of Reinforcement Learning (RL) post-training. As RL posttraining becomes pivotal for advancing LLM reasoning, the absence of specialized contamination detection methods in this paradigm presents a critical vulnerability. To address this, we conduct the first systematic study of data detection within RL post-training scenario and propose Self-Critique. Our method is motivated by a key observation: after RL phase, the output entropy distribution of LLMs tends to collapse into highly specific and sparse modes. Self-Critique probes for the underlying policy collapse, i.e., the model's convergence to a narrow reasoning path, which causes this entropy reduction. To facilitate this research, we also introduce RL-MIA, a benchmark constructed to simulate this specific contamination scenario. Extensive experiments show that Self-Critique significantly outperforms baseline methods across multiple models and contamination tasks, achieving an AUC improvement of up to 30%. Whereas existing methods are close to a random guess for RL-phase contamination, our method makes detection possible. Our benchmark and code are available at https://github.com/yongding-tao/RL-Data-Contamination .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2b8b4de-9eb8-428a-8da4-43f5c71e4b5eBuilds on26
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
Related papers
- ConStat: Performance-Based Contamination Detection in Large Language ModelsJasper Dekoninck, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 40 citations
- On The Fragility of Benchmark Contamination Detection in Reasoning ModelsHan Wang, Haoyu Li, Brian Ko, Huan ZhangICLR 2026 · 8 citations
- Min-K%++: Improved Baseline for Pre-Training Data Detection from Large Language ModelsJingyang Zhang, Jingwei Sun, Eric C. Yeats, Yang Ouyang et al.ICLR 2025
- From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMsGuangyu Shen, Siyuan Cheng, Xiangzhe Xu, Yuan Zhou et al.ICML 2026
- Probing Language Models for Pre-training Data DetectionZhenhua Liu, Tong Zhu, Chuanyuan Tan, Bing Liu et al.ACL 2024 · 4 citations
