An Exponential Improvement on the Memorization Capacity of Deep Threshold Networks
Shashank Rajput, Kartik Sreenivasan, Dimitris S. Papailiopoulos, Amin Karbasi
摘要
It is well known that modern deep neural networks are powerful enough to memorize datasets even when the labels have been randomized. Recently, Vershynin (2020) settled a long standing question by Baum (1988), proving that deep threshold networks can memorize points in dimensions using neurons and weights, where is the minimum distance between the points. In this work, we improve the dependence on from exponential to almost linear, proving that neurons and weights are sufficient. Our construction uses Gaussian random weights only in the first layer, while all the subsequent layers use binary or integer weights. We also prove new lower bounds by connecting memorization in neural networks to the purely geometric problem of separating points on a sphere using hyperplanes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- On the Optimal Memorization Power of ReLU Neural NetworksGal Vardi, Gilad Yehudai, Ohad ShamirICLR 2022 · 被引用 42 次
- Why Robust Generalization in Deep Learning is Difficult: Perspective of Expressive PowerBinghui Li, Jikai Jin, Han Zhong, John E. Hopcroft 等NeurIPS 2022 · 被引用 37 次
- Memorization Capacity of Multi-Head Attention in TransformersSadegh Mahdavi, Renjie Liao, Christos ThrampoulidisICLR 2024 · 被引用 34 次
- Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?Tokio Kajitsuka, Issei SatoICLR 2024 · 被引用 31 次
- Size and depth of monotone neural networks: interpolation and approximationDan Mikulincer, Daniel ReichmanNeurIPS 2022 · 被引用 14 次
它引用的顶会 Paper1
相关 Paper
- Network size and size of the weights in memorization with two-layers neural networksSébastien Bubeck, Ronen Eldan, Yin Tat Lee, Dan MikulincerNeurIPS 2020 · 被引用 28 次
- Neural Networks Learning and Memorization with (almost) no Over-ParameterizationAmit DanielyNeurIPS 2020 · 被引用 38 次
- Optimal robust Memorization with ReLU Neural NetworksLijia Yu, Xiao-Shan Gao, Lijun ZhangICLR 2024 · 被引用 4 次
- A Law of Data Reconstruction for Random Features (And Beyond)Leonardo Iurada, Simone Bombari, Tatiana Tommasi, Marco MondelliICLR 2026 · 被引用 3 次
- Generalizablity of Memorization Neural NetworkLijia Yu, Xiao-Shan Gao, Lijun Zhang, Yibo MiaoNeurIPS 2024 · 被引用 5 次
