Dataset Distillation for Memorized Data: Soft Labels can Leak Held-Out Teacher Knowledge
Freya Behrens, Lenka Zdeborová
摘要
Dataset distillation aims to compress training data into fewer examples via a teacher, from which a student can learn effectively. While its success is often attributed to structure in the data, modern neural networks also memorize specific facts, but if and how such memorized information can be transferred in distillation settings remains less understood. While this transfer may be desirable in some applications, it also raises privacy concerns, where preventing such leakage is crucial. In this work, we show that students trained on soft labels from teachers can indeed achieve non-trivial accuracy on held-out memorized data they never directly observed. This effect persists on structured data when the teacher has not generalized. To understand this effect in isolation, we consider finite random i.i.d. datasets where generalization is a priori impossible and a successful teacher fit implies pure memorization. Still, students can learn non-trivial information about the held-out data, in some cases up to perfect accuracy. For multinomial logistic classification and single layer MLPs, we show this corresponds to the setting where the teacher can be recovered functionally -- the student matches the teacher's predictions on all possible inputs, including the held-out memorized data. We empirically show that these phenomena strongly depend on the sample complexity and the temperature with which the logits are smoothed, but persist across varying network capacities, architectures and dataset compositions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Towards Understanding Subliminal Learning: When and How Hidden Biases TransferSimon Schrodi, Elias Kempf, Fazl Barez, Thomas BroxICLR 2026 · 被引用 30 次
- ManifoldGD: Training-Free Hierarchical Manifold Guidance for Diffusion-Based Dataset DistillationAyush Roy, Wei-Yang Alex Lee, Rudrasis Chakraborty, Vishnu Suresh LokhandeCVPR 2026 · 被引用 2 次
它引用的顶会 Paper10
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff PerspectiveHelong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou 等ICLR 2021 · 被引用 209 次
- A statistical perspective on distillationAditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim 等ICML 2021 · 被引用 97 次
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith 等ICLR 2023 · 被引用 54 次
- A Label is Worth A Thousand Images in Dataset DistillationTian Qin, Zhiwei Deng, David Alvarez-MelisNeurIPS 2024 · 被引用 39 次
- What is Dataset Distillation Learning?William Yang, Ye Zhu, Zhiwei Deng, Olga RussakovskyICML 2024 · 被引用 13 次
相关 Paper
- Students Parrot Their Teachers: Membership Inference on Model DistillationMatthew Jagielski, Milad Nasr, Katherine Lee, Christopher A. Choquette-Choo 等NeurIPS 2023 · 被引用 53 次
- Remember the Past: Distilling Datasets into Addressable Memories for Neural NetworksZhiwei Deng, Olga RussakovskyNeurIPS 2022 · 被引用 140 次
- Toward Understanding Adversarial Distillation: Why Robust Teachers FailHongsin Lee, Hye Won ChungICML 2026
- What Knowledge Gets Distilled in Knowledge Distillation?Utkarsh Ojha, Yuheng Li, Anirudh Sundara Rajan, Yingyu Liang 等NeurIPS 2023 · 被引用 58 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
