Lune

NeurIPS2022Top-tier venue

Most Activation Functions Can Win the Lottery Without Excessive Depth

Rebekka Burkholz

2022Year
27Citations
15Top-tier citations

Abstract

The strong lottery ticket hypothesis has highlighted the potential for training deep neural networks by pruning, which has inspired interesting practical and theoretical insights into how neural networks can represent functions. For networks with ReLU activation functions, it has been proven that a target network with depth LL can be approximated by the subnetwork of a randomly initialized neural network that has double the target's depth 2L2L and is wider by a logarithmic factor. We show that a depth L+1L+1 network is sufficient. This result indicates that we can expect to find lottery tickets at realistic, commonly used depths while only requiring logarithmic overparametrization. Our novel construction approach applies to a large class of activation functions and is not limited to ReLUs.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers15

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines