Lune

ICLR2021顶会

The inductive bias of ReLU networks on orthogonally separable data

Mary Phuong, Christoph H. Lampert

出版方
2021年份
53被引次数
33顶会引用

摘要

We study the inductive bias of two-layer ReLU networks trained by gradient flow. We identify a class of easy-to-learn (`orthogonally separable') datasets, and characterise the solution that ReLU networks trained on such datasets converge to. Irrespective of network width, the solution turns out to be a combination of two max-margin classifiers: one corresponding to the positive data subset and one corresponding to the negative data subset. The proof is based on the recently introduced concept of extremal sectors, for which we prove a number of properties in the context of orthogonal separability. In particular, we prove stationarity of activation patterns from some time TT onwards, which enables a reduction of the ReLU network to an ensemble of linear subnetworks.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper33

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖