Effect of Activation Functions on the Training of Overparametrized Neural Nets
Abhishek Panigrahi, Abhishek Shetty, Navin Goyal
Abstract
It is well-known that overparametrized neural networks trained using gradient-based methods quickly achieve small training error with appropriate hyperparameter settings. Recent papers have proved this statement theoretically for highly overparametrized networks under reasonable assumptions. These results either assume that the activation function is ReLU or they crucially depend on the minimum eigenvalue of a certain Gram matrix depending on the data, random initialization and the activation function. In the later case, existing works only prove that this minimum eigenvalue is non-zero and do not provide quantitative bounds. On the empirical side, a contemporary line of investigations has proposed a number of alternative activation functions which tend to perform better than ReLU at least in some settings but no clear understanding has emerged. This state of affairs underscores the importance of theoretically understanding the impact of activation functions on training. In the present paper, we provide theoretical results about the effect of activation function on the training of highly overparametrized 2-layer neural networks. A crucial property that governs the performance of an activation is whether or not it is smooth. For non-smooth activations such as ReLU, SELU and ELU, all eigenvalues of the associated Gram matrix are large under minimal assumptions on the data. For smooth activations such as tanh, swish and polynomials, the situation is more complex. If the subspace spanned by the data has small dimension then the minimum eigenvalue of the Gram matrix can be small leading to slow training. But if the dimension is large and the data satisfies another mild condition, then the eigenvalues are large. If we allow deep networks, then the small data dimension is not a limitation provided that the depth is sufficient. We discuss a number of extensions and applications of these results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8951d0c7-4a05-4690-b9e9-c97dc815be80Cited by top-tier papers7
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 327 citations
- Statistical-Query Lower Bounds via Functional GradientsSurbhi Goel, Aravind Gollakota, Adam R. KlivansNeurIPS 2020 · 72 citations
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear NetworkJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICML 2021 · 26 citations
- New Complexity-Theoretic Frontiers of Tractability for Neural Network TrainingCornelius Brand, Robert Ganian, Mathis RoctonNeurIPS 2023 · 4 citations
- DiGRAF: Diffeomorphic Graph-Adaptive Activation FunctionKrishna Sri Ipsit Mantri, Xinzhi Wang, Carola-Bibiane Schönlieb, Bruno Ribeiro et al.NeurIPS 2024 · 3 citations
Related papers
- Rational neural networksNicolas Boullé, Yuji Nakatsukasa, Alex TownsendNeurIPS 2020 · 130 citations
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello et al.ICML 2023 · 12 citations
- Low Curvature Activations Reduce Overfitting in Adversarial TrainingVasu Singla, Sahil Singla, Soheil Feizi, David JacobsICCV 2021 · 49 citations
- The Implicit Bias of Minima Stability: A View from Function SpaceRotem Mulayoff, Tomer Michaeli, Daniel SoudryNeurIPS 2021 · 65 citations
- Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal DataYiwen Kou, Zixiang Chen, Quanquan GuNeurIPS 2023 · 24 citations
