Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum Technique
Koushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar Pandey
摘要
Deep learning researchers have a keen interest in proposing new novel activation functions that can boost neural network performance. A good choice of activation function can have a significant effect on improving network performance and training dynamics. Rectified Linear Unit (ReLU) is a popular hand-designed activation function and is the most common choice in the deep learning community due to its simplicity though ReLU has some drawbacks. In this paper, we have proposed two new novel activation functions based on approximation of the maximum function, and we call these functions Smooth Maximum Unit (SMU and SMU-1). We show that SMU and SMU-1 can smoothly approximate ReLU, Leaky ReLU, or more general Maxout family, and GELU is a particular case of SMU. Replacing ReLU by SMU, Top-1 classification accuracy improves by 6.22%, 3.39%, 3.51%, and 3.08% on the CIFAR100 dataset with ShuffleNet V2, PreActResNet-50, ResNet-50, and SeNet-50 models respectively. Also, our experimental evaluation shows that SMU and SMU-1 improve network performance in a variety of deep learning tasks like image classification, object detection, semantic segmentation, and machine translation compared to widely used activation functions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Explaining Time Series via Contrastive and Locally Sparse PerturbationsZichuan Liu, Yingying Zhang, Tianchun Wang, Zefan Wang 等ICLR 2024 · 被引用 26 次
- The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant StepsizeDongyan Lucy Huo, Yixuan Zhang, Yudong Chen, Qiaomin XieNeurIPS 2024 · 被引用 9 次
- Koopman-based generalization bound: New aspect for full-rank weightsYuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda 等ICLR 2024 · 被引用 6 次
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 被引用 3 次
- IIEU: Rethinking Neural Feature Activation from Decision-MakingSudong CaiICCV 2023 · 被引用 1 次
相关 Paper
- ErfAct and Pserf: Non-monotonic Smooth Trainable Activation FunctionsKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyAAAI 2022 · 被引用 16 次
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 被引用 17 次
- Activate or Not: Learning Customized ActivationNingning Ma, Xiangyu Zhang, Ming Liu, Jian SunCVPR 2021
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 被引用 10 次
- Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning DynamicsIndrashis Das, Mahmoud Safari, Steven Adriaensen, Frank HutterNeurIPS 2025 · 被引用 2 次
