Activate or Not: Learning Customized Activation
Ningning Ma, Xiangyu Zhang, Ming Liu, Jian Sun
Abstract
We present a simple, effective, and general activation function we term ACON which learns to activate the neurons or not. Interestingly, we find Swish, the recent popular NAS-searched activation, can be interpreted as a smooth approximation to ReLU. Intuitively, in the same way, we approximate the more general Maxout family to our novel ACON family, which remarkably improves the performance and makes Swish a special case of ACON. Next, we present meta-ACON, which explicitly learns to optimize the parameter switching between non-linear (activate) and linear (inactivate) and provides a new design space. By simply changing the activation function, we show its effectiveness on both small models and highly optimized large models (e.g. it improves the ImageNet top-1 accuracy rate by 6.7% and 1.8% on respectively). Moreover, our novel ACON can be naturally transferred to object detection and semantic segmentation, showing that ACON is an effective alternative in a variety of tasks. Code is available at https: // github . com/ nmaac/ acon .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44d581c0-6faf-4c2b-8eba-ae69d17f8844Cited by top-tier papers8
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi et al.CVPR 2024 · 137 citations
- ErfAct and Pserf: Non-monotonic Smooth Trainable Activation FunctionsKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyAAAI 2022 · 16 citations
- Human-Inspired Facial Sketch Synthesis with Dynamic AdaptationFei Gao, Yifan Zhu, Chang Jiang, Nannan WangICCV 2023 · 10 citations
- Unsupervised Salient Instance DetectionXin Tian, Ke Xu, Rynson W. H. LauCVPR 2024 · 2 citations
- IIEU: Rethinking Neural Feature Activation from Decision-MakingSudong CaiICCV 2023 · 1 citation
Builds on2
Related papers
- Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum TechniqueKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyCVPR 2022 · 61 citations
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 17 citations
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 10 citations
- Stochastic Adaptive Activation FunctionKyungsu Lee, Jaeseung Yang, Haeyun Lee, Jae Youn HwangNeurIPS 2022 · 6 citations
- Efficient Activation Function Optimization through Surrogate ModelingGarrett Bingham, Risto MiikkulainenNeurIPS 2023 · 10 citations
