Activate or Not: Learning Customized Activation
Ningning Ma, Xiangyu Zhang, Ming Liu, Jian Sun
摘要
We present a simple, effective, and general activation function we term ACON which learns to activate the neurons or not. Interestingly, we find Swish, the recent popular NAS-searched activation, can be interpreted as a smooth approximation to ReLU. Intuitively, in the same way, we approximate the more general Maxout family to our novel ACON family, which remarkably improves the performance and makes Swish a special case of ACON. Next, we present meta-ACON, which explicitly learns to optimize the parameter switching between non-linear (activate) and linear (inactivate) and provides a new design space. By simply changing the activation function, we show its effectiveness on both small models and highly optimized large models (e.g. it improves the ImageNet top-1 accuracy rate by 6.7% and 1.8% on respectively). Moreover, our novel ACON can be naturally transferred to object detection and semantic segmentation, showing that ACON is an effective alternative in a variety of tasks. Code is available at https: // github . com/ nmaac/ acon .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image RestorationShihao Zhou, Duosheng Chen, Jinshan Pan, Jinglei Shi 等CVPR 2024 · 被引用 137 次
- ErfAct and Pserf: Non-monotonic Smooth Trainable Activation FunctionsKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyAAAI 2022 · 被引用 16 次
- Human-Inspired Facial Sketch Synthesis with Dynamic AdaptationFei Gao, Yifan Zhu, Chang Jiang, Nannan WangICCV 2023 · 被引用 10 次
- Unsupervised Salient Instance DetectionXin Tian, Ke Xu, Rynson W. H. LauCVPR 2024 · 被引用 2 次
- IIEU: Rethinking Neural Feature Activation from Decision-MakingSudong CaiICCV 2023 · 被引用 1 次
它引用的顶会 Paper2
相关 Paper
- Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum TechniqueKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyCVPR 2022 · 被引用 61 次
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 被引用 17 次
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 被引用 10 次
- Stochastic Adaptive Activation FunctionKyungsu Lee, Jaeseung Yang, Haeyun Lee, Jae Youn HwangNeurIPS 2022 · 被引用 6 次
- Efficient Activation Function Optimization through Surrogate ModelingGarrett Bingham, Risto MiikkulainenNeurIPS 2023 · 被引用 10 次
