Knowledge Distillation as Decontamination? Revisiting the "Data Laundering" Concern in Classification Tasks
Hengyu Luo, Raúl Vázquez, Timothee Mickus, Filip Ginter, Jörg Tiedemann
摘要
Concerns have been raised that knowledge distillation may transfer test-set knowledge from a contaminated teacher to a clean student-a "data laundering" effect that potentially threatens evaluation integrity. In this paper, we assess the severity of this phenomenon. If these concerns regarding data laundering are minor, then distillation could be used to mitigate risks of direct data exposure. Across eight classification benchmarks, we find that substantial laundering is the exception rather than the rule: unlike the large performance gains from direct contamination, any accuracy inflation from laundering is consistently smaller and statistically insignificant in all but two cases. More broadly, using sample-level analysis, we find that the two phenomena are weakly correlated, suggesting that laundering is not simply a diluted form of contamination but a distinct effect that arises primarily when benchmarks exhibit large train-test distribution gaps. Motivated by this, we conduct controlled experiments that systematically enlarge the train-test distance on two benchmarks where laundering was initially negligible, and observe that laundering becomes more significant as the gap widens. Taken together, our results indicate that knowledge distillation, despite rare benchmarkspecific residues, can be expected to function as an effective decontamination technique that largely mitigates test-data leakage. Code is available at https: //github.com/hengyu-luo/kd-revisiting-data-laundering-concern .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled CorpusJesse Dodge, Maarten Sap, Ana Marasovic, William Agnew 等EMNLP 2021 · 被引用 18 次
- Revisiting Data-Free Knowledge Distillation with Poisoned TeachersJunyuan Hong, Yi Zeng, Shuyang Yu, Lingjuan Lyu 等ICML 2023 · 被引用 16 次
- Data Laundering: Artificially Boosting Benchmark Results through Knowledge DistillationJonibek Mansurov, Akhmed Sakip, Alham Fikri AjiACL 2025 · 被引用 4 次
相关 Paper
- Analyzing the Confidentiality of Undistillable Teachers in Knowledge DistillationSouvik Kundu, Qirui Sun, Yao Fu, Massoud Pedram 等NeurIPS 2021 · 被引用 35 次
- Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine TranslationMuhammed Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch, Jiaming Luo 等ICML 2025
- How Contaminated Is Your Benchmark? Measuring Dataset Leakage in Large Language Models with Kernel DivergenceHyeong Kyu Choi, Maxim Khanov, Hongxin Wei, Yixuan LiICML 2025
- What Knowledge Gets Distilled in Knowledge Distillation?Utkarsh Ojha, Yuheng Li, Anirudh Sundara Rajan, Yingyu Liang 等NeurIPS 2023 · 被引用 58 次
- Adversarially Robust DistillationMicah Goldblum, Liam Fowl, Soheil Feizi, Tom GoldsteinAAAI 2020 · 被引用 258 次
