Lune

ICLR2021Top-tier venue

The inductive bias of ReLU networks on orthogonally separable data

Mary Phuong, Christoph H. Lampert

2021Year
53Citations
33Top-tier citations

Abstract

We study the inductive bias of two-layer ReLU networks trained by gradient flow. We identify a class of easy-to-learn (`orthogonally separable') datasets, and characterise the solution that ReLU networks trained on such datasets converge to. Irrespective of network width, the solution turns out to be a combination of two max-margin classifiers: one corresponding to the positive data subset and one corresponding to the negative data subset. The proof is based on the recently introduced concept of extremal sectors, for which we prove a number of properties in the context of orthogonal separability. In particular, we prove stationarity of activation patterns from some time TT onwards, which enables a reduction of the ReLU network to an ensemble of linear subnetworks.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 44be48e0-b4ec-4c5a-a308-5fe50478caea

Cited by top-tier papers33

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines