Padé Activation Units: End-to-end Learning of Flexible Activation Functions in Deep Networks
Alejandro Molina, Patrick Schramowski, Kristian Kersting
摘要
The performance of deep network learning strongly depends on the choice of the non-linear activation function associated with each neuron. However, deciding on the best activation is non-trivial, and the choice depends on the architecture, hyper-parameters, and even on the dataset. Typically these activations are fixed by hand before training. Here, we demonstrate how to eliminate the reliance on first picking fixed activation functions by using flexible parametric rational functions instead. The resulting Pade Activation Units (PAUs) can both approximate common activation functions and also learn new ones while providing compact representations. Our empirical evidence shows that end-to-end learning deep networks with PAUs can increase the predictive performance. Moreover, PAUs pave the way to approximations with provable robustness. this https URL
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Rational neural networksNicolas Boullé, Yuji Nakatsukasa, Alex TownsendNeurIPS 2020 · 被引用 130 次
- Adaptive Rational Activations to Boost Deep Reinforcement LearningQuentin Delfosse, Patrick Schramowski, Martin Mundt, Alejandro Molina 等ICLR 2024 · 被引用 25 次
- Learning specialized activation functions with the Piecewise Linear UnitYucong Zhou, Zezhou Zhu, Zhao ZhongICCV 2021 · 被引用 17 次
- ErfAct and Pserf: Non-monotonic Smooth Trainable Activation FunctionsKoushik Biswas, Sandeep Kumar, Shilpak Banerjee, Ashish Kumar PandeyAAAI 2022 · 被引用 16 次
- CoFrNets: Interpretable Neural Architecture Inspired by Continued FractionsIsha Puri, Amit Dhurandhar, Tejaswini Pedapati, Karthikeyan Shanmugam 等NeurIPS 2021 · 被引用 13 次
相关 Paper
- Differential Equation Units: Learning Functional Forms of Activation Functions from DataMohamadAli Torkamani, Shiv Shankar, Amirmohammad Rooshenas, Phillip WallisAAAI 2020 · 被引用 3 次
- Rational Neural Networks have Expressivity AdvantagesMaosen Tang, Alex TownsendICML 2026 · 被引用 1 次
- Fractional Adaptive Linear UnitsJulio Zamora, Anthony D. Rhodes, Lama NachmanAAAI 2022 · 被引用 10 次
- Deep Learning for Functional Data Analysis with Adaptive Basis LayersJunwen Yao, Jonas Mueller, Jane-Ling WangICML 2021 · 被引用 40 次
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 被引用 8 次
