ErfAct and Pserf: Non-monotonic Smooth Trainable Activation Functions
Koushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar Pandey
摘要
An activation function is a crucial component of a neural network that introduces non-linearity in the network. The state-of-the-art performance of a neural network depends also on the perfect choice of an activation function. We propose two novel non-monotonic smooth trainable activation functions, called ErfAct and Pserf. Experiments suggest that the proposed functions improve the network performance significantly compared to the widely used activations like ReLU, Swish, and Mish. Replacing ReLU by ErfAct and Pserf, we have 5.68% and 5.42% improvement for top-1 accuracy on Shufflenet V2 (2.0x) network in CIFAR100 dataset, 2.11% and 1.96% improvement for top-1 accuracy on Shufflenet V2 (2.0x) network in CIFAR10 dataset, 1.0%, and 1.0% improvement on mean average precision (mAP) on SSD300 model in Pascal VOC dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- IIEU: Rethinking Neural Feature Activation from Decision-MakingSudong CaiICCV 2023 · 被引用 1 次
- Toward Principled Flexible Scaling for Self-Gated Neural ActivationSudong Cai, Shuyuan Zheng, Bingzhi Chen, Shuai Yuan 等ICLR 2026
- AdaShift: Learning Discriminative Self-Gated Neural Feature Activation With an Adaptive Shift FactorSudong CaiCVPR 2024
它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep NetworksAlejandro Molina, Patrick Schramowski, Kristian KerstingICLR 2020 · 被引用 116 次
- Activate or Not: Learning Customized ActivationNingning Ma, Xiangyu Zhang, Ming Liu, Jian SunCVPR 2021
相关 Paper
- Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum TechniqueKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyCVPR 2022 · 被引用 61 次
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 被引用 17 次
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 被引用 10 次
- Effect of Activation Functions on the Training of Overparametrized Neural NetsAbhishek Panigrahi, Abhishek Shetty, Navin GoyalICLR 2020 · 被引用 24 次
- Rational Neural Networks have Expressivity AdvantagesMaosen Tang, Alex TownsendICML 2026 · 被引用 1 次
