Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum Technique
Koushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar Pandey
Abstract
Deep learning researchers have a keen interest in proposing new novel activation functions that can boost neural network performance. A good choice of activation function can have a significant effect on improving network performance and training dynamics. Rectified Linear Unit (ReLU) is a popular hand-designed activation function and is the most common choice in the deep learning community due to its simplicity though ReLU has some drawbacks. In this paper, we have proposed two new novel activation functions based on approximation of the maximum function, and we call these functions Smooth Maximum Unit (SMU and SMU-1). We show that SMU and SMU-1 can smoothly approximate ReLU, Leaky ReLU, or more general Maxout family, and GELU is a particular case of SMU. Replacing ReLU by SMU, Top-1 classification accuracy improves by 6.22%, 3.39%, 3.51%, and 3.08% on the CIFAR100 dataset with ShuffleNet V2, PreActResNet-50, ResNet-50, and SeNet-50 models respectively. Also, our experimental evaluation shows that SMU and SMU-1 improve network performance in a variety of deep learning tasks like image classification, object detection, semantic segmentation, and machine translation compared to widely used activation functions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af83ff5d-27cb-446e-b6bd-b8d3ef826d2aCited by top-tier papers9
- Explaining Time Series via Contrastive and Locally Sparse PerturbationsZichuan Liu, Yingying Zhang, Tianchun Wang, Zefan Wang et al.ICLR 2024 · 26 citations
- The Collusion of Memory and Nonlinearity in Stochastic Approximation With Constant StepsizeDongyan Lucy Huo, Yixuan Zhang, Yudong Chen, Qiaomin XieNeurIPS 2024 · 9 citations
- Koopman-based generalization bound: New aspect for full-rank weightsYuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda et al.ICLR 2024 · 6 citations
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 3 citations
- IIEU: Rethinking Neural Feature Activation from Decision-MakingSudong CaiICCV 2023 · 1 citation
Related papers
- ErfAct and Pserf: Non-monotonic Smooth Trainable Activation FunctionsKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyAAAI 2022 · 16 citations
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 17 citations
- Activate or Not: Learning Customized ActivationNingning Ma, Xiangyu Zhang, Ming Liu, Jian SunCVPR 2021
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 10 citations
- Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning DynamicsIndrashis Das, Mahmoud Safari, Steven Adriaensen, Frank HutterNeurIPS 2025 · 2 citations
