A Law of Data Reconstruction for Random Features (And Beyond)
Leonardo Iurada, Simone Bombari, Tatiana Tommasi, Marco Mondelli
摘要
Large-scale deep learning models are known to memorize parts of the training set. In machine learning theory, memorization is often framed as interpolation or label fitting, and classical results show that this can be achieved when the number of parameters in the model is larger than the number of training samples . In this work, we consider memorization from the perspective of data reconstruction, demonstrating that this can be achieved when is larger than , where is the dimensionality of the data. More specifically, we show that, in the random features model, when , the subspace spanned by the training samples in feature space gives sufficient information to identify the individual samples in input space. Our analysis suggests an optimization method to reconstruct the dataset from the model parameters, and we demonstrate that this method performs well on various architectures (random features, two-layer fully-connected and deep residual networks). Our results reveal a law of data reconstruction, according to which the entire training dataset can be recovered as exceeds the threshold .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
相关 Paper
- Generalizablity of Memorization Neural NetworkLijia Yu, Xiao-Shan Gao, Lijun Zhang, Yibo MiaoNeurIPS 2024 · 被引用 5 次
- Deconstructing Data Reconstruction: Multiclass, Weight Decay and General LossesGon Buzaglo, Niv Haim, Gilad Yehudai, Gal Vardi 等NeurIPS 2023 · 被引用 32 次
- Simulating Training Dynamics to Reconstruct Training Data from Deep Neural NetworksHanling Tian, Yuhang Liu, Mingzhen He, Zhengbao He 等ICLR 2025
- The Curious Case of Benign MemorizationSotiris Anagnostidis, Gregor Bachmann, Lorenzo Noci, Thomas HofmannICLR 2023 · 被引用 1 次
- Optimal robust Memorization with ReLU Neural NetworksLijia Yu, Xiao-Shan Gao, Lijun ZhangICLR 2024 · 被引用 4 次
