A Law of Data Reconstruction for Random Features (And Beyond)
Leonardo Iurada, Simone Bombari, Tatiana Tommasi, Marco Mondelli
Abstract
Large-scale deep learning models are known to memorize parts of the training set. In machine learning theory, memorization is often framed as interpolation or label fitting, and classical results show that this can be achieved when the number of parameters in the model is larger than the number of training samples . In this work, we consider memorization from the perspective of data reconstruction, demonstrating that this can be achieved when is larger than , where is the dimensionality of the data. More specifically, we show that, in the random features model, when , the subspace spanned by the training samples in feature space gives sufficient information to identify the individual samples in input space. Our analysis suggests an optimization method to reconstruct the dataset from the model parameters, and we demonstrate that this method performs well on various architectures (random features, two-layer fully-connected and deep residual networks). Our results reveal a law of data reconstruction, according to which the entire training dataset can be recovered as exceeds the threshold .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d8248c2-bee4-434f-8765-8c09f3bd5e84Builds on31
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 674 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 260 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
Related papers
- Generalizablity of Memorization Neural NetworkLijia Yu, Xiao-Shan Gao, Lijun Zhang, Yibo MiaoNeurIPS 2024 · 5 citations
- Deconstructing Data Reconstruction: Multiclass, Weight Decay and General LossesGon Buzaglo, Niv Haim, Gilad Yehudai, Gal Vardi et al.NeurIPS 2023 · 32 citations
- Simulating Training Dynamics to Reconstruct Training Data from Deep Neural NetworksHanling Tian, Yuhang Liu, Mingzhen He, Zhengbao He et al.ICLR 2025
- The Curious Case of Benign MemorizationSotiris Anagnostidis, Gregor Bachmann, Lorenzo Noci, Thomas HofmannICLR 2023 · 1 citation
- Optimal robust Memorization with ReLU Neural NetworksLijia Yu, Xiao-Shan Gao, Lijun ZhangICLR 2024 · 4 citations
