Efficient Activation Function Optimization through Surrogate Modeling
Garrett Bingham, Risto Miikkulainen
Abstract
Carefully designed activation functions can improve the performance of neural networks in many machine learning tasks. However, it is difficult for humans to construct optimal activation functions, and current activation function search algorithms are prohibitively expensive. This paper aims to improve the state of the art through three steps: First, the benchmark datasets Act-Bench-CNN, Act-Bench-ResNet, and Act-Bench-ViT were created by training convolutional, residual, and vision transformer architectures from scratch with 2,913 systematically generated activation functions. Second, a characterization of the benchmark space was developed, leading to a new surrogate-based method for optimization. More specifically, the spectrum of the Fisher information matrix associated with the model's predictive distribution at initialization and the activation function's output distribution were found to be highly predictive of performance. Third, the surrogate was used to discover improved activation functions in several real-world tasks, with a surprising finding: a sigmoidal design that outperformed all other activation functions was discovered, challenging the status quo of always using rectifier nonlinearities in deep learning. Each of these steps is a contribution in its own right; together they serve as a practical and theoretical foundation for further research on activation function optimization. Introduction Activation functions are an important choice in neural network design [2, 46] . In order to realize the benefits of good activation functions, researchers often design new functions based on characteristics like smoothness, groundedness, monotonicity, and limit behavior. While these properties have proven useful, humans are ultimately limited by design biases and by the relatively small number of functions they can consider. On the other hand, automated search methods can evaluate thousands of unique functions, and as a result, often discover better activation functions than those designed by humans. However, such approaches do not usually have a theoretical justification, and instead focus only on performance. This limitation results in computationally inefficient ad hoc algorithms that may miss good solutions and may not scale to large models and datasets. This paper addresses these drawbacks in a data-driven way through three steps. First, in order to provide a foundation for theory and algorithm development, convolutional, residual, and vision transformer based architectures were trained from scratch with 2,913 different activation functions, resulting in three activation function benchmark datasets: Act-Bench-CNN, Act-Bench-ResNet, * GB is currently a research scientist at Google DeepMind. AQuaSurF code is available at https:// github.com/cognizant-ai-labs/aquasurf , and the benchmark datasets are at https://github.com/ cognizant-ai-labs/act-bench . 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10abf277-87bf-4e0b-953b-e3b126269266Cited by top-tier papers1
Ask how each one uses itBuilds on9
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- CoAtNet: Marrying Convolution and Attention for All Data SizesZihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing TanNeurIPS 2021 · 1,747 citations
- Neural Architecture Search without TrainingJoe Mellor, Jack Turner, Amos Storkey, Elliot J. CrowleyICML 2021 · 477 citations
- How Powerful are Performance Predictors in Neural Architecture Search?Colin White, Arber Zela, Robin Ru, Yang Liu et al.NeurIPS 2021 · 168 citations
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep NetworksAlejandro Molina, Patrick Schramowski, Kristian KerstingICLR 2020 · 116 citations
Related papers
- AutoInit: Analytic Signal-Preserving Weight Initialization for Neural NetworksGarrett Bingham, Risto MiikkulainenAAAI 2023 · 6 citations
- Activate or Not: Learning Customized ActivationNingning Ma, Xiangyu Zhang, Ming Liu, Jian SunCVPR 2021
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 17 citations
- Generalization Properties of NAS under Activation and Skip Connection SearchZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2022 · 23 citations
- Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the WildDamien Teney, Liangze Jiang, Florin Gogianu, Ehsan AbbasnejadCVPR 2025
