Learning specialized activation functions with the Piecewise Linear Unit
Yucong Zhou, Zezhou Zhu, Zhao Zhong
摘要
The choice of activation functions is crucial for modern deep neural networks. Popular hand-designed activation functions like Rectified Linear Unit(ReLU) and its variants show promising performance in various tasks and models. Swish, the automatically discovered activation function, has been proposed and outperforms ReLU on many challenging datasets. However, it has two main drawbacks. First, the tree-based search space is highly discrete and restricted, which is difficult for searching. Second, the sample-based searching method is inefficient, making it infeasible to find specialized activation functions for each dataset or neural architecture. To tackle these drawbacks, we propose a new activation function called Piecewise Linear Unit(PWLU), which incorporates a carefully designed formulation and learning method. It can learn specialized activation functions and achieves SOTA performance on large-scale datasets like ImageNet and COCO. For example, on ImageNet classification dataset, PWLU improves 0.9%/0.53%/1.0%/1.7%/1.0% top-1 accuracy over Swish for ResNet-18/ResNet-50/MobileNet-V2/MobileNetV3/EfficientNet-B0. PWLU is also easy to implement and efficient at inference, which can be widely applied in real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Collaboration of Experts: Achieving 80% Top-1 Accuracy on ImageNet with 100M FLOPsYikang Zhang, Zhuo Chen, Zhao ZhongICML 2022 · 被引用 11 次
- IIEU: Rethinking Neural Feature Activation from Decision-MakingSudong CaiICCV 2023 · 被引用 1 次
- PPLNs: Parametric Piecewise Linear Networks for Event-Based Temporal Modeling and BeyondChen Song, Zhenxiao Liang, Bo Sun, Qixing HuangNeurIPS 2024 · 被引用 1 次
- AdaShift: Learning Discriminative Self-Gated Neural Feature Activation With an Adaptive Shift FactorSudong CaiCVPR 2024
它引用的顶会 Paper3
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep NetworksAlejandro Molina, Patrick Schramowski, Kristian KerstingICLR 2020 · 被引用 116 次
相关 Paper
- Smooth Maximum Unit: Smooth Activation Function for Deep Networks using Smoothing Maximum TechniqueKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyCVPR 2022 · 被引用 61 次
- Activate or Not: Learning Customized ActivationNingning Ma, Xiangyu Zhang, Ming Liu, Jian SunCVPR 2021
- ErfAct and Pserf: Non-monotonic Smooth Trainable Activation FunctionsKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyAAAI 2022 · 被引用 16 次
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 被引用 10 次
- Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning DynamicsIndrashis Das, Mahmoud Safari, Steven Adriaensen, Frank HutterNeurIPS 2025 · 被引用 2 次
