Lune

NeurIPS2020Top-tier venue

Network size and size of the weights in memorization with two-layers neural networks

Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, Dan Mikulincer

2020Year
28Citations
13Top-tier citations

Abstract

In 1988, Eric B. Baum showed that two-layers neural networks with threshold activation function can perfectly memorize the binary labels of n points in general position in R d using only n/d neurons. We observe that with ReLU networks, using four times as many neurons one can fit arbitrary real labels. Moreover, for approximate memorization up to error ε, the neural tangent kernel can also memorize with only O n d • log(1/ε) neurons (assuming that the data is well dispersed too). We show however that these constructions give rise to networks where the magnitude of the neurons' weights are far from optimal. In contrast we propose a new training procedure for ReLU networks, based on complex (as opposed to real) recombination of the neurons, for which we show approximate memorization with both O n d • log(1/ε) ε neurons, as well as nearly-optimal size of the weights.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bb07ad7d-58da-413e-9d16-6e314a0c5aca

Cited by top-tier papers13

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines